arXiv ScienceSearch

arXiv subjects

Zichen Song

Publications and source records attributed to Zichen Song.

16 recordsLinked to original sources

AGI and the Limits of Value Production

This paper develops a political-economy model of artificial general intelligence (AGI) as a technology that progressively substitutes living labor with machine-based productive systems. The model studies the transition from the first moment at which AGI becomes economically capable of replacing labor to the later moment at which AGI becomes technically and actually capable of near-complete replacement. The central distinction is between technical substitutability and actual adoption. Technical substitutability is the feasible replacement ceiling implied by the state of AGI capability, whereas actual adoption is the realized replacement share chosen under cost, profitability, and adoption frictions. Under the strict value-theoretic assumption that AGI transfers value but does not itself create new value, deeper AGI adoption raises the organic composition of capital, reduces the quantity of living labor when adoption outpaces the creation of new labor fields, compresses the source of surplus value, and places downward pressure on the social rate of profit. In the limiting case in which actual AGI adoption approaches complete substitution and new labor fields fail to compensate for displaced labor, living labor tends to zero, surplus value tends to zero, and the profit rate tends to zero. The model therefore identifies near-complete AGI substitution not merely as an efficiency transition, but as a boundary case for value production under a strict political-economy theory of value.

cs.CE

PPCR-IM: A System for Multi-layer DAG-based Public Policy Consequence Reasoning and Social Indicator Mapping

Public policy decisions are typically justified using a narrow set of headline indicators, leaving many downstream social impacts unstructured and difficult to compare across policies. We propose PPCR-IM, a system for multi-layer DAG-based consequence reasoning and social indicator mapping that addresses this gap. Given a policy description and its context, PPCR-IM uses an LLM-driven, layer-wise generator to construct a directed acyclic graph of intermediate consequences, allowing child nodes to have multiple parents to capture joint influences. A mapping module then aligns these nodes to a fixed indicator set and assigns one of three qualitative impact directions: increase, decrease, or ambiguous change. For each policy episode, the system outputs a structured record containing the DAG, indicator mappings, and three evaluation measures: an expected-indicator coverage score, a discovery rate for overlooked but relevant indicators, and a relative focus ratio comparing the systems coverage to that of the government. PPCR-IM is available both as an online demo and as a configurable XLSX-to-JSON batch pipeline.

cs.SI

qNEP: A highly efficient neuroevolution potential with dynamic charges for large-scale atomistic simulations

Although electrostatics can be incorporated into machine-learned interatomic potentials, existing approaches are computationally very demanding, limiting large-scale, long-time simulations of electrostatics-driven phenomena such as dielectric response, infrared activity, and field-matter coupling. Here, we extend the neuroevolution potential (NEP), a highly efficient machine-learned interatomic potential, to a charge-aware framework (qNEP) by introducing explicit, environment-dependent partial charges. Each ionic partial charge is represented by a neural network as a function of the local descriptor vector, analogous to the NEP site-energy model. This formulation enables the direct prediction of the Born effective charge tensor for each ion and, consequently, the polarization. As a result, dielectric properties, infrared spectra, and coupling to external electric fields can be evaluated within a unified framework. We derive consistent expressions for the forces and virials that explicitly account for the position dependence of the partial charges. The qNEP method has been implemented in the free-and-open-source GPUMD package, with support for both Ewald summation and particle-particle particle-mesh treatments of electrostatics. We demonstrate the accuracy and efficiency of the qNEP approach through representative applications to water, Li7La3Zr2O12, BaTiO3, and a magnesium-water interface. These results show that qNEP enables accurate atomistic simulations with explicit long-range electrostatics, scalable to million-atom systems on nanosecond time scales using consumer-grade GPUs.

physics.comp-ph

The multilinear fractional bounded mean oscillation operator theory I: sparse domination, sparse $T1$ theorem, off-diagonal extrapolation, quantitative weighted estimate -- for generalized commutators

This paper introduces and studies a class of multilinear fractional bounded mean oscillation operators (denoted {\rm $m$-FBMOOs}) defined on ball-basis measure spaces $(X, \mu, \mathcal{B})$. These operators serve as a generalization of canonical classes, such as the multilinear fractional maximal operators, the multilinear fractional Ahlfors-Beurling operators, the multilinear pseudo-differential operators with multi-parameter H\"ormander symbol, and some multilinear operators admitting $\mathbb{V}$-valued $m$-linear fractional Dini-type Calder\'on-Zygmund kernel representation. Crucially, the definition utilized here, incorporating the notion of "bounded mean oscillation," provides greater generality compared to those in Karagulyan (2019) and Cao et al. (2023). Our investigation systematically examines the properties of these operators and their generalized commutators through the lens of modern harmonic analysis, focusing on two principal directions: 1.We establish Karagulyan-type sparse domination for the generalized commutators. Subsequently, by developing bespoke dyadic representation theorems pertinent to this setting, we prove a corresponding multilinear fractional sparse $T1$ theorem for these generalized commutators. 2.With sparse bounds established, we obtain weighted estimates in multiple complementary methods: (1) Under a novel class of multilinear fractional weights, we prove four types of weighted inequalities: sharp-type, Bloom-type, mixed weak-type, and local decay-type.(2) We develop a multilinear non-diagonal extrapolation framework for these weights, which transfers boundedness flexibly among weighted spaces and establishes corresponding vector-valued inequalities.

math.CA

The multilinear fractional sparse operator theory I: pointwise domination and weighted estimate

How to establish some specific quantitative weighted estimates for the generalized commutator of multilinear fractional singular integral operator $\mathcal{T}_{\eta}^{{\bf b}}$ is the focus of this paper, which is defined by $$\mathcal{T}_{\eta}^{{\bf b}}(\vec{f})(x):= \mathcal{T}_{\eta}\left((b_1(x) - b_1)^{\beta_1}f_1,\ldots,(b_m(x) - b_m)^{\beta_m}f_m\right)(x),$$ where $\mathcal{T}_{\eta}$ is a multilinear fractional singular integral operator, ${\bf b}:=({b_1}, \cdots ,{b_m})$ is a set of symbol functions, and $({\beta_1}, \cdots ,{\beta_m}) \in {\mathbb{N}_0^m}$. Pointwise dominating the aforementioned commutator leads us to consider a class of higher order multi-symbol multilinear fractional sparse operator ${\mathcal A}_{\eta ,\mathcal{S},\tau}^\mathbf{b,k,t}$ to achieve this long-cherished wish. Therefore, it suffices to construct its quantitative weighted estimates, which firstly include the characterization of several types of multilinear weighted conditions $A_{\vec p,q}^*$, $W_{\vec p,q}^\infty$, and $H_{\vec p,q}^\infty$. Within the scope of this work, Bloom type estimate for first order multi-symbol multilinear fractional sparse operator is established herein. Moreover, we derive two distinct Bloom type estimates for higher order multi-symbol multilinear fractional sparse operator by using "maximal weight method" and "iterated weight method" respectively, which not only refines some of Lerner's methods but greatly enhances the generality of our conclusions. Endpoint quantitative estimates for multilinear fractional singular integral operators and their first order commutators are also obtained as the last main result. It is also worthy of highlighting that some important multilinear fractional operators are applicable to our results as applications.

math.CA

EM-MIAs: Enhancing Membership Inference Attacks in Large Language Models through Ensemble Modeling

With the widespread application of large language models (LLM), concerns about the privacy leakage of model training data have increasingly become a focus. Membership Inference Attacks (MIAs) have emerged as a critical tool for evaluating the privacy risks associated with these models. Although existing attack methods, such as LOSS, Reference-based, min-k, and zlib, perform well in certain scenarios, their effectiveness on large pre-trained language models often approaches random guessing, particularly in the context of large-scale datasets and single-epoch training. To address this issue, this paper proposes a novel ensemble attack method that integrates several existing MIAs techniques (LOSS, Reference-based, min-k, zlib) into an XGBoost-based model to enhance overall attack performance (EM-MIAs). Experimental results demonstrate that the ensemble model significantly improves both AUC-ROC and accuracy compared to individual attack methods across various large language models and datasets. This indicates that by combining the strengths of different methods, we can more effectively identify members of the model's training data, thereby providing a more robust tool for evaluating the privacy risks of LLM. This study offers new directions for further research in the field of LLM privacy protection and underscores the necessity of developing more powerful privacy auditing methods.

cs.RO

Layer Importance and Hallucination Analysis in Large Language Models via Enhanced Activation Variance-Sparsity

Evaluating the importance of different layers in large language models (LLMs) is crucial for optimizing model performance and interpretability. This paper first explores layer importance using the Activation Variance-Sparsity Score (AVSS), which combines normalized activation variance and sparsity to quantify each layer's contribution to overall model performance. By ranking layers based on AVSS and pruning the least impactful 25\%, our experiments on tasks such as question answering, language modeling, and sentiment classification show that over 90\% of the original performance is retained, highlighting potential redundancies in LLM architectures. Building on AVSS, we propose an enhanced version tailored to assess hallucination propensity across layers (EAVSS). This improved approach introduces Hallucination-Specific Activation Variance (HSAV) and Hallucination-Specific Sparsity (HSS) metrics, allowing precise identification of hallucination-prone layers. By incorporating contrastive learning on these layers, we effectively mitigate hallucination generation, contributing to more robust and efficient LLMs(The maximum performance improvement is 12\%). Our results on the NQ, SciQ, TriviaQA, TruthfulQA, and WikiQA datasets demonstrate the efficacy of this method, offering a comprehensive framework for both layer importance evaluation and hallucination mitigation in LLMs.

cs.CL

AVSS: Layer Importance Evaluation in Large Language Models via Activation Variance-Sparsity Analysis

The evaluation of layer importance in deep learning has been an active area of research, with significant implications for model optimization and interpretability. Recently, large language models (LLMs) have gained prominence across various domains, yet limited studies have explored the functional importance and performance contributions of individual layers within LLMs, especially from the perspective of activation distribution. In this work, we propose the Activation Variance-Sparsity Score (AVSS), a novel metric combining normalized activation variance and sparsity to assess each layer's contribution to model performance. By identifying and removing approximately the lowest 25% of layers based on AVSS, we achieve over 90% of original model performance across tasks such as question answering, language modeling, and sentiment classification, indicating that these layers may be non-essential. Our approach provides a systematic method for identifying less critical layers, contributing to efficient large language model architectures.

cs.CL

Multimodal growth and development assessment model

With the development of social economy and the improvement of people's attention to health, the growth and development of children and adolescents has become an important indicator to measure the level of national health. Therefore, accurate and timely assessment of children's growth and development has become increasingly important. At the same time, global health inequalities, especially child malnutrition and stunting in developing countries, urgently require effective assessment tools to monitor and intervene. In recent years, the rapid development of technologies such as big data, artificial intelligence, and cloud computing, and the cross-integration of multiple disciplines such as biomedicine, statistics, and computer science have promoted the rapid development of large-scale models for growth and development assessment. However, there are still problems such as too single evaluation factors, inaccurate diagnostic results, and inability to give accurate and reasonable recommendations. The multi-modal growth and development assessment model uses the public data set of RSNA ( North American College of Radiology ) as the training set, and the data set of the Department of Pediatrics of Huaibei People's Hospital as the open source test set. The embedded ICL module enables the model to quickly adapt and identify the tasks that need to be done to ensure that under the premise of considering multiple evaluation factors, accurate diagnosis results and reasonable medical recommendations are given, so as to provide solutions to the above problems and promote the development of the medical field.

cs.CE

CCSRP: Robust Pruning of Spiking Neural Networks through Cooperative Coevolution

Spiking neural networks (SNNs) have shown promise in various dynamic visual tasks, yet those ready for practical deployment often lack the compactness and robustness essential in resource-limited and safety-critical settings. Prior research has predominantly concentrated on enhancing the compactness or robustness of artificial neural networks through strategies like network pruning and adversarial training, with little exploration into similar methodologies for SNNs. Robust pruning of SNNs aims to reduce computational overhead while preserving both accuracy and robustness. Current robust pruning approaches generally necessitate expert knowledge and iterative experimentation to establish suitable pruning criteria or auxiliary modules, thus constraining their broader application. Concurrently, evolutionary algorithms (EAs) have been employed to automate the pruning of artificial neural networks, delivering remarkable outcomes yet overlooking the aspect of robustness. In this work, we propose CCSRP, an innovative robust pruning method for SNNs, underpinned by cooperative co-evolution. Robust pruning is articulated as a tri-objective optimization challenge, striving to balance accuracy, robustness, and compactness concurrently, resolved through a cooperative co-evolutionary pruning framework that independently prunes filters across layers using EAs. Our experiments on CIFAR-10 and SVHN demonstrate that CCSRP can match or exceed the performance of the latest methodologies.

cs.NE

Case-based reasoning approach for diagnostic screening of children with developmental delays

According to the World Health Organization, the population of children with developmental delays constitutes approximately 6% to 9% of the total population. Based on the number of newborns in Huaibei, Anhui Province, China, in 2023 (94,420), it is estimated that there are about 7,500 cases (suspected cases of developmental delays) of suspicious cases annually. Early identification and appropriate early intervention for these children can significantly reduce the wastage of medical resources and societal costs. International research indicates that the optimal period for intervention in children with developmental delays is before the age of six, with the golden treatment period being before three and a half years of age. Studies have shown that children with developmental delays who receive early intervention exhibit significant improvement in symptoms; some may even fully recover. This research adopts a hybrid model combining a CNN-Transformer model with Case-Based Reasoning (CBR) to enhance the screening efficiency for children with developmental delays. The CNN-Transformer model is an excellent model for image feature extraction and recognition, effectively identifying features in bone age images to determine bone age. CBR is a technique for solving problems based on similar cases; it solves current problems based on past experiences, similar to how humans solve problems through learning from experience. Given CBR's memory capability to judge and compare new cases based on previously stored old cases, it is suitable for application in support systems with latent and variable characteristics. Therefore, this study utilizes the CNN-Transformer-CBR to establish a screening system for children with developmental delays, aiming to improve screening efficiency.

cs.CV

FTS: A Framework to Find a Faithful TimeSieve

The field of time series forecasting has garnered significant attention in recent years, prompting the development of advanced models like TimeSieve, which demonstrates impressive performance. However, an analysis reveals certain unfaithfulness issues, including high sensitivity to random seeds and minute input noise perturbations. Recognizing these challenges, we embark on a quest to define the concept of \textbf{\underline{F}aithful \underline{T}ime\underline{S}ieve \underline{(FTS)}}, a model that consistently delivers reliable and robust predictions. To address these issues, we propose a novel framework aimed at identifying and rectifying unfaithfulness in TimeSieve. Our framework is designed to enhance the model's stability and resilience, ensuring that its outputs are less susceptible to the aforementioned factors. Experimentation validates the effectiveness of our proposed framework, demonstrating improved faithfulness in the model's behavior. Looking forward, we plan to expand our experimental scope to further validate and optimize our algorithm, ensuring comprehensive faithfulness across a wide range of scenarios. Ultimately, we aspire to make this framework can be applied to enhance the faithfulness of not just TimeSieve but also other state-of-the-art temporal methods, thereby contributing to the reliability and robustness of temporal modeling as a whole.

cs.LG

Extrapolation to product Morrey-Herz spaces and applications

The purpose of this paper is threefold. First, we introduce product Morrey-Herz spaces and product block-Herz spaces, establish their duality, and prove the boundedness of the strong maximal operator on product block-Herz spaces; these results provide the foundation for extrapolation. Second, using the Rubio de Francia iteration method, we establish extrapolation results on product Morrey-Herz spaces. Finally, we give applications to Fefferman-Stein vector-valued strong maximal inequalities, the John-Nirenberg inequality, a characterization of little bmo in terms of product Morrey-Herz spaces, and the boundedness of bi-parameter Calder\'on-Zygmund operators and their commutators.

math.FA

New fractional type weights and the boundedness of some operators

Two classes of fractional type variable weights are established in this paper. The first kind of weights ${A_{\vec p( \cdot ),q( \cdot )}}$ are variable multiple weights, which are characterized by the weighted variable boundedness of multilinear fractional type operators, called multilinear Hardy--Littlewood--Sobolev theorem on weighted variable Lebesgue spaces. Meanwhile, the weighted variable boundedness for the commutators of multilinear fractional type operators are also obtained. This generalizes some known work, such as Moen (2009), Bernardis--Dalmasso--Pradolini (2014), and Cruz-Uribe--Guzm\'an (2020). Another class of weights ${{\mathbb{A}}_{p( \cdot ),q(\cdot)}}$ are variable matrix weights that also characterized by certain fractional type operators. This generalize some previous results on matrix weights ${{\mathbb{A}}_{p( \cdot )}}$.

math.CA

A Comprehensive Review of Community Detection in Graphs

The study of complex networks has significantly advanced our understanding of community structures which serves as a crucial feature of real-world graphs. Detecting communities in graphs is a challenging problem with applications in sociology, biology, and computer science. Despite the efforts of an interdisciplinary community of scientists, a satisfactory solution to this problem has not yet been achieved. This review article delves into the topic of community detection in graphs, which serves as a thorough exposition of various community detection methods from perspectives of modularity-based method, spectral clustering, probabilistic modelling, and deep learning. Along with the methods, a new community detection method designed by us is also presented. Additionally, the performance of these methods on the datasets with and without ground truth is compared. In conclusion, this comprehensive review provides a deep understanding of community detection in graphs.

cs.SI

Shared and Private Information Learning in Multimodal Sentiment Analysis with Deep Modal Alignment and Self-supervised Multi-Task Learning

Designing an effective representation learning method for multimodal sentiment analysis tasks is a crucial research direction. The challenge lies in learning both shared and private information in a complete modal representation, which is difficult with uniform multimodal labels and a raw feature fusion approach. In this work, we propose a deep modal shared information learning module based on the covariance matrix to capture the shared information between modalities. Additionally, we use a label generation module based on a self-supervised learning strategy to capture the private information of the modalities. Our module is plug-and-play in multimodal tasks, and by changing the parameterization, it can adjust the information exchange relationship between the modes and learn the private or shared information between the specified modes. We also employ a multi-task learning strategy to help the model focus its attention on the modal differentiation training data. We provide a detailed formulation derivation and feasibility proof for the design of the deep modal shared information learning module. We conduct extensive experiments on three common multimodal sentiment analysis baseline datasets, and the experimental results validate the reliability of our model. Furthermore, we explore more combinatorial techniques for the use of the module. Our approach outperforms current state-of-the-art methods on most of the metrics of the three public datasets.

cs.CL