arXiv ScienceSearch

arXiv subjects

Ying Song

Publications and source records attributed to Ying Song.

At least 19 recordsLinked to original sources

Optical-Phonon-Enabled Large Lattice Thermal Conductivity Anisotropy in Hexagonal Perovskites Cs$BX_3$ ($B$ = Mg, Cd; $X$ = Cl, Br, I)

Materials exhibiting strongly anisotropic lattice thermal conductivity are desirable for thermal-management applications, yet such behavior is commonly associated with layered or quasi-one-dimensional van der Waals crystals and highly anisotropic elastic properties. Here, we investigate lattice thermal transport in the hexagonal perovskites Cs$BX_3$ ($B=$ Mg, Cd; $X=$ Cl, Br, I) using first-principles calculations. At 300~K, the calculated in-plane and out-of-plane lattice thermal conductivities range from 0.13--0.83 and 0.34--6.26~Wm$^{-1}$K$^{-1}$, respectively, corresponding to anisotropy ratios of 2.6--7.5. This pronounced anisotropy is remarkable given the relatively modest elastic anisotropy, characterized by $C_{33}/C_{11}$ = 0.994--1.842. Our analysis reveals that medium-frequency optical phonons provide an efficient out-of-plane heat-transport channel, contrary to the conventional picture in which heat transport is dominated by acoustic phonons. These findings identify face-sharing octahedral frameworks as a promising platform for engineering strong thermal-conductivity anisotropy in mechanically near-isotropic, non--van der Waals crystals.

cond-mat.mtrl-sci

Large Language Models for Fuzz Testing in Microservices: A Systematic Literature Review

Microservice systems (MSS) increasingly rely on heterogeneous APIs whose combinatorial input space and stateful dependencies challenge traditional fuzz testing. Meanwhile, Large Language Models (LLMs) have recently been introduced to enhance fuzzing with semantic reasoning over specifications, inputs, and runtime feedback. This paper presents a systematic literature review (SLR) of LLM-assisted fuzz testing for microservices to synthesise how LLMs are applied, evaluated, and what challenges remain. Following established SLR guidelines, we analyze 20 primary studies published between 2024 and 2026. Results show LLMs are mainly used as semantic input generators in black-box fuzzing, with a growing shift towards agent-based and retrieval-augmented architectures, improving valid input generation and modestly increasing coverage and vulnerability detection. However, evaluation remains heterogeneous, with limited benchmark standardization, scarce cost reporting, and a bias toward single-service experiments, highlighting a gap with real-world multi-service systems. This review provides a taxonomy of LLM roles and integrations, a consolidated view of evaluation practices, and a mapping of open challenges to research directions, supporting the design and deployment of LLM-driven fuzzing in microservices.

cs.SE

Towards LLM Accelerated Rapid Reviews for Software Tool Discovery -- Case for Log Anomaly Detection

In software engineering research, the primary outcome is frequently a tool. However, for practitioners and academics alike, it is hard to tell which tools are maintained and do they work out of the box. In this paper, we propose a pipeline to identify relevant studies with LLM screening, extract the tools presented in them, and run them with LLM-based coding agent. To evaluate the feasibility of our approach we focus on software log anomaly detection tools. We begin the study by designing a broad search string that yields 3233 hits from Scopus. We request two LLMs to provide an inclusion probability for each title-abstract pair according to the inclusion and exclusion criteria. From the 3233 exported abstracts, this screening reduced the number of included papers to 569, out of which we could download 470. These papers included 206 unique links and after manual evaluation we determined 83 to be tools. Finally, we ran the LLM-based coding agent on these 83 links, and got 24 successfully running tools. As replicating our approach would require roughly only 4 hours of human effort, of which 3 hours were manual PDF downloading, and 12 hours of LLM running time, this demonstrates promising efficiency when utilizing LLMs in rapid reviews. Because practitioner-built tools often lack academic papers, in the future we aim to expand our analysis to tool-hosting platforms such as GitHub and PyPI. In the future, we plan to formalize our workflow as LLM Agent Skills to make our approach easier to adopt.

cs.SE

How Do Ice Shelves Calve? Peridynamic Modeling of Ice Shelf Fracture Driven by Wave Erosion, Basal Melting, and Buoyancy Flexure

An ice shelf is a floating extension of a land-based ice sheet into the ocean. It plays a crucial role in slowing down the flow of land ice into the sea, thus stabilizing the ice sheet. However, this stabilizing effect can be weakened by ice calving, a process in which large fragments of ice detach from the ice shelf. Although ice calving is widely acknowledged as a major contributor to ice mass loss, and its frequency and magnitude are highly sensitive to the environmental forcing, the underlying physics-based mechanisms remain poorly understood, particularly under ocean wave actions. In this context, we developed a nonlocal peridynamics (PD) framework to model the ice calving process subjected to wave-induced frontal corrosion. The proposed physics-based PD framework enables investigation of the coupled effects of self-weight bending, buoyancy-induced foot loosening, and ice calving process. To authors' best knowledge, this work represents the first attempt to employ a physics-based peridynamics framework for simulating ice calving processes. Compared with conventional finite element methods (FEM), the PD framework naturally captures crack initiation, interaction, and propagation without the need for special numerical treatments, thereby providing a robust tool for simulating fracture phenomena under large deformations and long-term environmental loading. To quantitatively resolve fracture processes, we implemented a static first Piola Kirchhoff virial stress formulation within the PD framework, allowing direct evaluation of stress concentration and energy release at evolving crack tips. Subsequently, the model is rigorously validated through one-to-one comparisons with finite-element stress fields, analytical beam-theory solutions, and recent field observations of wave-driven ice-shelf failure reported by Sartore et al. (2025).

cs.CE

A Comparative Study of Semantic Log Representations for Software Log-based Anomaly Detection

Recent deep learning (DL) methods for log anomaly detection increasingly rely on semantic log representation methods that convert the textual content of log events into vector embeddings as input to DL models. However, these DL methods are typically evaluated as end-to-end pipelines, while the impact of different semantic representation methods is not well understood. In this paper, we benchmark widely used semantic log representation methods, including static word embedding methods (Word2Vec, GloVe, and FastText) and the BERT-based contextual embedding method, across diverse DL models for log-event level anomaly detection on three publicly available log datasets: BGL, Thunderbird, and Spirit. We identify an effectiveness--efficiency trade off under CPU deployment settings: the BERT-based method is more effective, but incurs substantially longer log embedding generation time, limiting its practicality; static word embedding methods are efficient but are generally less effective and may yield insufficient detection performance. Motivated by this finding, we propose QTyBERT, a novel semantic log representation method that better balances this trade-off. QTyBERT uses SysBE, a lightweight BERT variant with system-specific quantization, to efficiently encode log events into vector embeddings on CPUs, and leverages CroSysEh to enhance the semantic expressiveness of these log embeddings. CroSysEh is trained unsupervisedly using unlabeled logs from multiple systems to capture the underlying semantic structure of the BERT model's embedding space. We evaluate QTyBERT against existing semantic log representation methods. Our results show that, for the DL models, using QTyBERT-generated log embeddings achieves detection effectiveness comparable to or better than BERT-generated log embeddings, while bringing log embedding generation time closer to that of static word embedding methods.

cs.SE

Hidden Chiral Ferroelectricity in AgNbO$_3$ Perovskite

AgNbO$_3$ is a lead-free perovskite with considerable potential for energy storage and optoelectronic applications, yet its low-temperature crystal structure has remained controversial. In this Letter, we revisit its low-energy structural landscape using a systematic first-principles structural search based on symmetry-adapted phonon-mode theory. We uncover a previously unreported chiral ferroelectric phase with space group $R3$, which exhibits a large spontaneous polarization and a low polarization switching barrier, enabling polarization reversal under electric fields. Crucially, the structural chirality of this phase is intrinsically locked to the ferroelectric polarization, allowing electrical control of the chiral handedness. Consequently, chiral optical responses--including circular dichroism, circular photogalvanic effect, optical activity, and second-order nonlinear optics--can be reversibly switched by an external electric field. These results not only clarify the complex low-temperature structural behavior of AgNbO$_3$ but also establish a rare purely inorganic platform for electric-field-tunable chirality, opening a pathway toward ultrafast, electrically controlled chiral optoelectronics.

cond-mat.mtrl-sci

Taipan: A Query-free Transfer-based Multiple Sensitive Attribute Inference Attack Solely from Publicly Released Graphs

Graph-structured data underpin a wide spectrum of modern applications. However, complex graph topologies and homophilic patterns can facilitate attribute inference attacks (AIAs) by enabling sensitive information leakage to propagate across local neighborhoods. Existing AIAs predominantly assume that adversaries can probe sensitive attributes through repeated model queries. Such assumptions are often impractical in real-world settings due to stringent data protection regulations, prohibitive query budgets, and heightened detection risks, especially when inferring multiple sensitive attributes. More critically, this model-centric perspective obscures a pervasive blind spot: \textbf{intrinsic multiple sensitive information leakage arising solely from publicly released graphs.} To exploit this unexplored vulnerability, we introduce a new attack paradigm and propose \textbf{Taipan, the first query-free transfer-based attack framework for multiple sensitive attribute inference attacks on graphs (G-MSAIAs).} Taipan integrates \emph{Hierarchical Attack Knowledge Routing} to capture intricate inter-attribute correlations, and \emph{Prompt-guided Attack Prototype Refinement} to mitigate negative transfer and performance degradation. We further present a systematic evaluation framework tailored to G-MSAIAs. Extensive experiments on diverse real-world graph datasets demonstrate that Taipan consistently achieves strong attack performance across same-distribution settings and heterogeneous similar- and out-of-distribution settings with mismatched feature dimensionalities, and remains effective even under rigorous differential privacy guarantees. Our findings underscore the urgent need for more robust multi-attribute privacy-preserving graph publishing methods and data-sharing practices.

cs.CR

AnoMod: A Dataset for Anomaly Detection and Root Cause Analysis in Microservice Systems

Microservice systems (MSS) have become a predominant architectural style for cloud services. Yet the community still lacks high-quality, publicly available datasets for anomaly detection (AD) and root cause analysis (RCA) in MSS. Most benchmarks emphasize performance-related faults and provide only one or two monitoring modalities, limiting research on broader failure modes and cross-modal methods. To address these gaps, we introduce a new multimodal anomaly dataset built on two open-source microservice systems: SocialNetwork and TrainTicket. We design and inject four categories of anomalies (Ano): performance-level, service-level, database-level, and code-level, to emulate realistic anomaly modes. For each scenario, we collect five modalities (Mod): logs, metrics, distributed traces, API responses, and code coverage reports, offering a richer, end-to-end view of system state and inter-service interactions. We name our dataset, reflecting its unique properties, as AnoMod. This dataset enables (1) evaluation of cross-modal anomaly detection and fusion/ablation strategies, and (2) fine-grained RCA studies across service and code regions, supporting end-to-end troubleshooting pipelines that jointly consider detection and localization.

cs.SE

CuriGS: Curriculum-Guided Gaussian Splatting for Sparse View Synthesis

3D Gaussian Splatting (3DGS) has recently emerged as an efficient, high-fidelity representation for real-time scene reconstruction and rendering. However, extending 3DGS to sparse-view settings remains challenging because of supervision scarcity and overfitting caused by limited viewpoint coverage. In this paper, we present CuriGS, a curriculum-guided framework for sparse-view 3D reconstruction using 3DGS. CuriGS addresses the core challenge of sparse-view synthesis by introducing student views: pseudo-views sampled around ground-truth poses (teacher). For each teacher, we generate multiple groups of student views with different perturbation levels. During training, we follow a curriculum schedule that gradually unlocks higher perturbation level, randomly sampling candidate students from the active level to assist training. Each sampled student is regularized via depth-correlation and co-regularization, and evaluated using a multi-signal metric that combines SSIM, LPIPS, and an image-quality measure. For every teacher and perturbation level, we periodically retain the best-performing students and promote those that satisfy a predefined quality threshold to the training set, resulting in a stable augmentation of sparse training views. Experimental results show that CuriGS outperforms state-of-the-art baselines in both rendering fidelity and geometric consistency across various synthetic and real sparse-view scenes. Project page: https://zijian1026.github.io/CuriGS/

cs.CV

GraphToxin: Reconstructing Full Unlearned Graphs from Graph Unlearning

Graph unlearning has emerged as a promising solution to comply with "the right to be forgotten" regulations by enabling the removal of sensitive information upon request. However, this solution is not foolproof. The involvement of multiple parties creates new attack surfaces, and residual traces of deleted data can still remain in the unlearned graph neural networks (GNNs). These vulnerabilities can be exploited by attackers to recover the supposedly erased samples, thereby undermining the intended functionality of graph unlearning. In this work, we propose GraphToxin, the first full graph reconstruction attack against graph unlearning. Specifically, we introduce a novel curvature matching module to provide fine-grained guidance for unlearned graph recovery. We demonstrate that GraphToxin can successfully subvert the regulatory guarantees expected from graph unlearning, it can recover not only a deleted individual's information and personal links but also sensitive content from their connections, thereby posing substantially more detrimental threats. Furthermore, we extend GraphToxin to multiple-node removal under both white-box and black-box settings, showcasing its practical feasibility and potential to cause considerable harm. We highlight the necessity of worst-case analysis and propose a systematic evaluation framework to assess attack performance under both random and worst-case node removal scenarios. Our extensive experiments demonstrate the effectiveness and flexibility of GraphToxin. Notably, existing defense mechanisms are largely ineffective against this attack or even amplify its performance in some cases. Given the severe privacy risks posed by GraphToxin, our work underscores the urgent need for more effective and robust defenses.

cs.LG

Class Incremental Medical Image Segmentation via Prototype-Guided Calibration and Dual-Aligned Distillation

Class incremental medical image segmentation (CIMIS) aims to preserve knowledge of previously learned classes while learning new ones without relying on old-class labels. However, existing methods 1) either adopt one-size-fits-all strategies that treat all spatial regions and feature channels equally, which may hinder the preservation of accurate old knowledge, 2) or focus solely on aligning local prototypes with global ones for old classes while overlooking their local representations in new data, leading to knowledge degradation. To mitigate the above issues, we propose Prototype-Guided Calibration Distillation (PGCD) and Dual-Aligned Prototype Distillation (DAPD) for CIMIS in this paper. Specifically, PGCD exploits prototype-to-feature similarity to calibrate class-specific distillation intensity in different spatial regions, effectively reinforcing reliable old knowledge and suppressing misleading information from old classes. Complementarily, DAPD aligns the local prototypes of old classes extracted from the current model with both global prototypes and local prototypes, further enhancing segmentation performance on old categories. Comprehensive evaluations on two widely used multi-organ segmentation benchmarks demonstrate that our method outperforms state-of-the-art methods, highlighting its robustness and generalization capabilities.

cs.CV

Unbiased Platform-Level Causal Estimation for Search Systems: A Competitive Isolation PSM-DID Framework

Evaluating platform-level interventions in search-based two-sided marketplaces is fundamentally challenged by systemic effects such as spillovers and network interference. While widely used for causal inference, the PSM (Propensity Score Matching) - DID (Difference-in-Differences) framework remains susceptible to selection bias and cross-unit interference from unaccounted spillovers. In this paper, we introduced Competitive Isolation PSM-DID, a novel causal framework that integrates propensity score matching with competitive isolation to enable platform-level effect measurement (e.g., order volume, GMV) instead of item-level metrics in search systems. Our approach provides theoretically guaranteed unbiased estimation under mutual exclusion conditions, with an open dataset released to support reproducible research on marketplace interference (github.com/xxxx). Extensive experiments demonstrate significant reductions in interference effects and estimation variance compared to baseline methods. Successful deployment in a large-scale marketplace confirms the framework's practical utility for platform-level causal inference.

cs.AI

MotionFlow:Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation

Generating videos guided by camera trajectories poses significant challenges in achieving consistency and generalizability, particularly when both camera and object motions are present. Existing approaches often attempt to learn these motions separately, which may lead to confusion regarding the relative motion between the camera and the objects. To address this challenge, we propose a novel approach that integrates both camera and object motions by converting them into the motion of corresponding pixels. Utilizing a stable diffusion network, we effectively learn reference motion maps in relation to the specified camera trajectory. These maps, along with an extracted semantic object prior, are then fed into an image-to-video network to generate the desired video that can accurately follow the designated camera trajectory while maintaining consistent object motions. Extensive experiments verify that our model outperforms SOTA methods by a large margin.

cs.CV

Exploiting Unlabeled Structures through Task Consistency Training for Versatile Medical Image Segmentation

Versatile medical image segmentation (VMIS) targets the segmentation of multiple classes, while obtaining full annotations for all classes is often impractical due to the time and labor required. Leveraging partially labeled datasets (PLDs) presents a promising alternative; however, current VMIS approaches face significant class imbalance due to the unequal category distribution in PLDs. Existing methods attempt to address this by generating pseudo-full labels. Nevertheless, these typically require additional models and often result in potential performance degradation from label noise. In this work, we introduce a Task Consistency Training (TCT) framework to address class imbalance without requiring extra models. TCT includes a backbone network with a main segmentation head (MSH) for multi-channel predictions and multiple auxiliary task heads (ATHs) for task-specific predictions. By enforcing a consistency constraint between the MSH and ATH predictions, TCT effectively utilizes unlabeled anatomical structures. To avoid error propagation from low-consistency, potentially noisy data, we propose a filtering strategy to exclude such data. Additionally, we introduce a unified auxiliary uncertainty-weighted loss (UAUWL) to mitigate segmentation quality declines caused by the dominance of specific tasks. Extensive experiments on eight abdominal datasets from diverse clinical sites demonstrate our approach's effectiveness.

cs.CV

Harmonic generation of graphene quantum dots in Hartree-Fock approximation

We theoretically investigate harmonic generation in graphene quantum dots under linearly polarized optical pulses, focusing on excitonic effects. Combining the tight-binding model and the single-particle density matrix approach, we derive a semiconductor Bloch equation under a static-screened Hartree-Fock approximation. This framework characterizes the electron-electron interaction through local Hartree potentials for direct Coulomb interaction and nonlocal Fock potentials for exchange interaction. Distinct confgurations of Hartree and Fock terms yield various approximation methods, including independent-particle approximation, mean-feld approximation, random phase approximation, and excitonic effects. We thoroughly analyze how these approximation methods affect the electronic energy levels, linear optical absorption, and nonlinear harmonic generation. Within excitonic effects, we present the dependence of harmonic generation on the geometric variations of graphene quantum dots (sizes, triangular/hexagonal shapes, and armchair/zigzag edges) and the amplitude and polarization of electric fields. Our findings show that excitonic effects significantly enhance optical responses of graphene nanostructures. For a dot ensemble formed by randomly oriented graphene quantum dots, only odd-order harmonics exist along the polarization direction of the incident light. Crucially, harmonic generation in graphene quantum dots exhibits high tunability via geometric configuration, making them promising candidates for nonlinear optical nanodevices.

cond-mat.mes-hall

Crater-shaped Enrichment of $\mathrm{V}_\mathrm{Si}$ Color Centers in $4H$-SiC using Single-Pulse Near-Infrared Femtosecond Laser Processing

Currently, Si vacancy ($\mathrm{V}_\mathrm{Si}$) color centers in SiC are of significant interest due to their potential applications in quantum sensing and quantum communication. Meanwhile, the qualities of laser-induced color centers are well guaranteed. Femtosecond laser processing suffices for increasing the yield of $\mathrm{V}_\mathrm{Si}$ color centers in bulk materials and forms crater-shaped enriched regions on the surface. However, there is a notable absence of existing simulation methods to explain the mechanisms behind laser-assisted $\mathrm{V}_\mathrm{Si}$ color center generation. In this work, we design a three-dimensional molecular dynamics (3D-MD) model using an integral hemi-ellipsoidal shell mathematical model to simulate the interaction of Gaussian laser beams with bulk materials. Furthermore, we calculate the transmittance, absorption coefficient, refractive index, and reflectivity of $4H$-SiC. Then, the absorptance of a 1030 nm laser in 350 {\mu}m-thick $4H$-SiC material is abtained to simulate the energy loss during the actual processing. Finally, the study analyzes the movement trajectories of $\mathrm{V}_\mathrm{Si}$ color centers and explains the source of $\mathrm{V}_\mathrm{Si}$ on the surface. This analysis explains the reasons for the enrichment of color centers in the crater-shaped regions formed after laser deposition. Our work provides an effective 3D-MD modeling approach to study the processing mechanisms of laser interaction with semiconductor materials, offering insights into efficient $\mathrm{V}_\mathrm{Si}$ color center creation processes.

physics.optics

Krait: A Backdoor Attack Against Graph Prompt Tuning

Graph prompt tuning has emerged as a promising paradigm to effectively transfer general graph knowledge from pre-trained models to various downstream tasks, particularly in few-shot contexts. However, its susceptibility to backdoor attacks, where adversaries insert triggers to manipulate outcomes, raises a critical concern. We conduct the first study to investigate such vulnerability, revealing that backdoors can disguise benign graph prompts, thus evading detection. We introduce Krait, a novel graph prompt backdoor. Specifically, we propose a simple yet effective model-agnostic metric called label non-uniformity homophily to select poisoned candidates, significantly reducing computational complexity. To accommodate diverse attack scenarios and advanced attack types, we design three customizable trigger generation methods to craft prompts as triggers. We propose a novel centroid similarity-based loss function to optimize prompt tuning for attack effectiveness and stealthiness. Experiments on four real-world graphs demonstrate that Krait can efficiently embed triggers to merely 0.15% to 2% of training nodes, achieving high attack success rates without sacrificing clean accuracy. Notably, in one-to-one and all-to-one attacks, Krait can achieve 100% attack success rates by poisoning as few as 2 and 22 nodes, respectively. Our experiments further show that Krait remains potent across different transfer cases, attack types, and graph neural network backbones. Additionally, Krait can be successfully extended to the black-box setting, posing more severe threats. Finally, we analyze why Krait can evade both classical and state-of-the-art defenses, and provide practical insights for detecting and mitigating this class of attacks.

cs.LG

MAPPING: Debiasing Graph Neural Networks for Fair Node Classification with Limited Sensitive Information Leakage

Despite remarkable success in diverse web-based applications, Graph Neural Networks(GNNs) inherit and further exacerbate historical discrimination and social stereotypes, which critically hinder their deployments in high-stake domains such as online clinical diagnosis, financial crediting, etc. However, current fairness research that primarily craft on i.i.d data, cannot be trivially replicated to non-i.i.d. graph structures with topological dependence among samples. Existing fair graph learning typically favors pairwise constraints to achieve fairness but fails to cast off dimensional limitations and generalize them into multiple sensitive attributes; besides, most studies focus on in-processing techniques to enforce and calibrate fairness, constructing a model-agnostic debiasing GNN framework at the pre-processing stage to prevent downstream misuses and improve training reliability is still largely under-explored. Furthermore, previous work on GNNs tend to enhance either fairness or privacy individually but few probe into their interplays. In this paper, we propose a novel model-agnostic debiasing framework named MAPPING (\underline{M}asking \underline{A}nd \underline{P}runing and Message-\underline{P}assing train\underline{ING}) for fair node classification, in which we adopt the distance covariance($dCov$)-based fairness constraints to simultaneously reduce feature and topology biases in arbitrary dimensions, and combine them with adversarial debiasing to confine the risks of attribute inference attacks. Experiments on real-world datasets with different GNN variants demonstrate the effectiveness and flexibility of MAPPING. Our results show that MAPPING can achieve better trade-offs between utility and fairness, and mitigate privacy risks of sensitive information leakage.

cs.LG