arXiv ScienceSearch

arXiv subjects

Lin Hu

Publications and source records attributed to Lin Hu.

At least 19 recordsLinked to original sources

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.

cs.CL

Nonvolatile photoswitching of a Mott state via reversible stacking rearrangement

Nonvolatile control of the Mott transition is a central goal in correlated-electron physics, offering access to fascinating emergent states and great potential for technological applications. Compared to chemical or mechanical approaches, ultrafast optical excitation further promises a path to create and manipulate novel non-equilibrium phases with ultimate spatiotemporal precision. However, achieving a truly nonvolatile electronic phase transition in laser-excited Mott systems remains an elusive challenge. Here, we present a highly robust and reversible method for optical control of the Mott state in van der Waals systems. Specifically, using angle-resolved photoemission spectroscopy, we observe a nonvolatile Mott-to-metallic transition in the ultrafast laser-excited charge density wave (CDW) material 1T-TaSe2. Complementary theoretical calculations reveal that this transition originates from a rearrangement of the interlayer CDW stacking. This new stacking order, formed following the ultrafast quenching of the CDW, circumvents the need for large-scale atomic sliding. Intriguingly, it introduces a significant in-plane component to the electron hopping and effectively reduces the ratio of on-site Coulomb interaction to bandwidth, thereby suppressing the Mott state and stabilizing a metallic phase. Our results establish optical-control of interlayer stacking as a versatile strategy for inducing nonvolatile phase transitions, opening a new route to tailor correlated electronic phases and realize reconfigurable high-frequency devices.

cond-mat.str-el

JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis

Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models are predominately constrained to dyadic, turn-based interactions. To address this, we introduce JoyVoice, a novel anthropomorphic foundation model designed for flexible, boundary-free synthesis of up to eight speakers. Unlike conventional cascaded systems, JoyVoice employs a unified E2E-Transformer-DiT architecture that utilizes autoregressive hidden representations directly for diffusion inputs, enabling holistic end-to-end optimization. We further propose a MM-Tokenizer operating at a low bitrate of 12.5 Hz, which integrates multitask semantic and MMSE losses to effectively model both semantic and acoustic information. Additionally, the model incorporates robust text front-end processing via large-scale data perturbation. Experiments show that JoyVoice achieves state-of-the-art results in multilingual generation (Chinese, English, Japanese, Korean) and zero-shot voice cloning. JoyVoice achieves top-tier results on both the Seed-TTS-Eval Benchmark and multi-speaker long-form conversational voice cloning tasks, demonstrating superior audio quality and generalization. It achieves significant improvements in prosodic continuity for long-form speech, rhythm richness in multi-speaker conversations, paralinguistic naturalness, besides superior intelligibility. We encourage readers to listen to the demo at https://jea-speech.github.io/JoyVoice

cs.SD

Balanced 3DGS: Gaussian-wise Parallelism Rendering with Fine-Grained Tiling

3D Gaussian Splatting (3DGS) is increasingly attracting attention in both academia and industry owing to its superior visual quality and rendering speed. However, training a 3DGS model remains a time-intensive task, especially in load imbalance scenarios where workload diversity among pixels and Gaussian spheres causes poor renderCUDA kernel performance. We introduce Balanced 3DGS, a Gaussian-wise parallelism rendering with fine-grained tiling approach in 3DGS training process, perfectly solving load-imbalance issues. First, we innovatively introduce the inter-block dynamic workload distribution technique to map workloads to Streaming Multiprocessor(SM) resources within a single GPU dynamically, which constitutes the foundation of load balancing. Second, we are the first to propose the Gaussian-wise parallel rendering technique to significantly reduce workload divergence inside a warp, which serves as a critical component in addressing load imbalance. Based on the above two methods, we further creatively put forward the fine-grained combined load balancing technique to uniformly distribute workload across all SMs, which boosts the forward renderCUDA kernel performance by up to 7.52x. Besides, we present a self-adaptive render kernel selection strategy during the 3DGS training process based on different load-balance situations, which effectively improves training efficiency.

cs.CV

All-Electrical Layer-Spintronics in Altermagnetic Bilayer

Electrical manipulation of spin-polarized current is highly desirable yet tremendously challenging in developing ultracompact spintronic device technology. Here we propose a scheme to realize the all-electrical manipulation of spin-polarized current in an altermagnetic bilayer. Such a bilayer system can host layer-spin locking, in which one layer hosts a spin-polarized current while the other layer hosts a current with opposite spin polarization. An out-of-plane electric field breaks the layer degeneracy, leading to a gate-tunable spin-polarized current whose polarization can be fully reversed upon flipping the polarity of the electric field. Using first-principles calculations, we show that CrS bilayer with C-type antiferromagnetic exchange interaction exhibits a hidden layer-spin locking mechanism that enables the spin polarization of the transport current to be electrically manipulated via the layer degree of freedom. We demonstrate that sign-reversible spin polarization as high as 87% can be achieved at room temperature. This work presents the pioneering concept of layer-spintronics which synergizes altermagnetism and bilayer stacking to achieve efficient electrical control of spin.

cond-mat.mes-hall

Mapping Information in Feature Extraction Transformation for Chirp Signal

Chirp signals have established diverse applications caused by the capable of producing time-dependent linear frequencies. Most feature extraction transformation methods for chirp signals focus on enhancing the performance of transform methods but neglecting the information derived from the transformation process. Consequently, they may fail to fully exploit the information from observations, resulting in decreased performance under conditions of low signal-to-noise ratio and limited observations. In this work, we develop a novel post-processing method called mapping information model to addressing this challenge. The model establishes a link between the observation space and feature space in feature extraction transform, enabling interference suppression and obtain more accurate information by iteratively resampling and assigning weights in both spaces. Analysis of the iteration process reveals a continual increase in weight of signal samples and a gradual stability in weight of noise samples. The demonstration of the noise suppression in the iteration process and feature enhancement supports the effectiveness of the mapping information model. Furthermore, numerical simulations also affirm the high efficiency of the proposed model by showcasing enhanced signal detection and estimation performances without requiring additional observations. This superior model allows amplifying performance within feature extraction transformation for chirp signal processing under low SNR and limited observation conditions, opens up new opportunities for areas such as communication, biomedicine, and remote sensing.

eess.SP

Interlayer magnetic interactions and ferroelectricity in $\pi$/3-twisted CrX$_2$ (X = Se, Te) bilayers

Recently, two-dimensional (2D) bilayer magnetic systems have been widely studied. Their interlayer magnetic interactions play a vital role in the magnetic properties. In this paper, we theoretically studied the interlayer magnetic interactions, magnetic states and ferroelectricity of $\pi$/3-twisted CrX$_2$ (X = Se, Te) bilayers ($\pi$/3-CrX$_2$). Our study reveals that the lateral shift could switch the magnetic state of the $\pi$/3-CrSe$_2$ between interlayer ferromagnetic and antiferromagnetic, while just tuning the strength of the interlayer antiferromagnetic interactions in $\pi$/3-CrTe$_2$. Furthermore, the lateral shift can alter the off-plane electric polarization in both $\pi$/3-CrSe$_2$ and $\pi$/3-CrTe$_2$. These results show that stacking is an effective way to tune both the magnetic and ferroelectric properties of 1T-CrX$_2$ bilayers, making the 1T-CrX$_2$ bilayers hold promise for 2D spintronic devices.

cond-mat.mtrl-sci

RingMo-lite: A Remote Sensing Multi-task Lightweight Network with CNN-Transformer Hybrid Framework

In recent years, remote sensing (RS) vision foundation models such as RingMo have emerged and achieved excellent performance in various downstream tasks. However, the high demand for computing resources limits the application of these models on edge devices. It is necessary to design a more lightweight foundation model to support on-orbit RS image interpretation. Existing methods face challenges in achieving lightweight solutions while retaining generalization in RS image interpretation. This is due to the complex high and low-frequency spectral components in RS images, which make traditional single CNN or Vision Transformer methods unsuitable for the task. Therefore, this paper proposes RingMo-lite, an RS multi-task lightweight network with a CNN-Transformer hybrid framework, which effectively exploits the frequency-domain properties of RS to optimize the interpretation process. It is combined by the Transformer module as a low-pass filter to extract global features of RS images through a dual-branch structure, and the CNN module as a stacked high-pass filter to extract fine-grained details effectively. Furthermore, in the pretraining stage, the designed frequency-domain masked image modeling (FD-MIM) combines each image patch's high-frequency and low-frequency characteristics, effectively capturing the latent feature representation in RS data. As shown in Fig. 1, compared with RingMo, the proposed RingMo-lite reduces the parameters over 60% in various RS image interpretation tasks, the average accuracy drops by less than 2% in most of the scenes and achieves SOTA performance compared to models of the similar size. In addition, our work will be integrated into the MindSpore computing platform in the near future.

cs.CV

Misspecified Model Estimation and Its Impact on Predictions

We study a linear statistical model where outcomes depend on regressors with fixed population coefficients and observation-specific latent coefficients, along with measurement errors. A decision-maker estimates population coefficients and uses the estimates to predict the latent coefficients for a given observation. We analyze how misspecification of some population coefficients distorts predictions, investigating comparative statics with respect to: (1) residual information in regressors associated with misspecified coefficients after projecting out those associated with free coefficients, (2) alignment between misspecification vector and latent-to-coefficient mapping. Applications include employee rating with unconscious bias and LLM-mediated consumer research.

econ.TH

Confidential Signal Cancellation Phenomenon in Interference Alignment Networks: Cause and Cure

This paper investigates physical layer security (PLS) in wireless interference networks. Specifically, we consider confidential transmission from a legitimate transmitter (Alice) to a legitimate receiver (Bob), in the presence of non-colluding passive eavesdroppers (Eves), as well as multiple legitimate transceivers. To mitigate interference at legitimate receivers and enhance PLS, artificial noise (AN) aided interference alignment (IA) is explored. However, the conventional leakage minimization (LM) based IA may exhibit confidential signal cancellation phenomenon. We theoretically analyze the cause and then establish a condition under which this phenomenon will occur almost surely. Moreover, we propose a means of avoiding this phenomenon by integrating the max-eigenmode beamforming (MEB) into the traditional LM based IA. By assuming that only statistical channel state informations (CSIs) of Eves and local CSIs of legitimate users are available, we derive a closed form expression for the secrecy outage probability (SOP), and establish a condition under which positive secrecy rate is achievable. To enhance security performance, an SOP constrained secrecy rate maximization (SRM) problem is formulated and an efficient numerical method is developed for the optimal solution. Numerical results confirm the effectiveness and the usefulness of the proposed approach.

cs.IT

Rationally Inattentive Echo Chambers

Rationally inattentive players allocate limited attention capacities to biased primary sources and to other players as secondary sources to acquire information about an uncertain state. The resulting Poisson attention network stochastically transmits information from primary sources to a recipient either directly or indirectly through the other players. We give conditions for the rise of echo-chamber equilibria in which players restrict attention to their own-biased primary source and same-type peers. We characterize the peer attention networks within echo chambers and develop tools for their comparative statics. Our results explain why modern information environments foster echo-chamber formation, how small differences in attention capacities can magnify into large disparities in allocations, and why regulatory interventions such as altering user visibility on social media or mandating exposure to opposing views may backfire with unintended consequences.

econ.TH

Atomistic Mechanism Underlying the Si(111)-(7\times7) Surface Reconstruction Revealed by Artificial Neural-network Potential

The 7\times7 reconstruction of the Si(111) surface represents arguably the most fascinating surface reconstruction so far observed in nature. Yet, the atomistic mechanism underpinning its formation remains unclear after it was discovered sixty years ago. Experimentally, it is observed post priori so that analysis of its formation mechanism can only be carried out in analogy with archaeology. Theoretically, density-functional-theory (DFT) correctly predicts the Si(111)-(7\times7) ground state but is impractical to simulate its formation process; while empirical potentials failed to produce it as the ground state. Developing an artificial neural-network potential of DFT quality, we carried out accurate large-scale simulations to unravel the formation of the Si(111)-(7\times7) surface. We reveal a possible step-mediated atom-pop rate-limiting process that triggers massive non-conserved atomic rearrangements, most remarkably, a critical process of collective vacancy diffusion that mediates a sequence of selective dimer, corner-hole, stacking fault and dimer-line pattern formation, to fulfill the 7\times7 reconstruction. Our findings may not only solve the long-standing mystery of this famous surface reconstruction but also illustrate the power of machine learning in studying complex structures.

cond-mat.mtrl-sci

Electoral Accountability and Selection with Personalized Information Aggregation

We study a model of electoral accountability and selection whereby heterogeneous voters aggregate incumbent politician's performance data into personalized signals through paying limited attention. Extreme voters' signals exhibit an own-party bias, which hampers their ability to discern the good and bad performances of the incumbent. While this effect alone would undermine electoral accountability and selection, there is a countervailing effect stemming from partisan disagreement, which makes the centrist voter more likely to be pivotal. In case the latter's unbiased signal is very informative about the incumbent's performance, the combined effect on electoral accountability and selection can actually be a positive one. For this reason, factors that carry a negative connotation in every political discourse -- such as increasing mass polarization and shrinking attention span -- have ambiguous accountability and selection effects in general. Correlating voters' signals, if done appropriately, unambiguously improves electoral accountability and selection and, hence, voter welfare.

econ.TH

The Politics of Personalized News Aggregation

We study how personalized news aggregation for rationally inattentive voters (NARI) affects policy polarization and public opinion. In a two-candidate electoral competition model, an attention-maximizing infomediary aggregates source data about candidates' valence into easy-to-digest news. Voters decide whether to consume news, trading off the expected gain from improved expressive voting against the attention cost. NARI generates policy polarization even if candidates are office-motivated. Personalized news aggregation makes extreme voters the disciplining entity of policy polarization, and the skewness of their signals is crucial for sustaining a high degree of policy polarization in equilibrium. Analysis of disciplining voters yields insights into the equilibrium and welfare consequences of regulating infomediaries.

econ.GN

GSI: GPU-friendly Subgraph Isomorphism

Subgraph isomorphism is a well-known NP-hard problem that is widely used in many applications, such as social network analysis and query over the knowledge graph. Due to the inherent hardness, its performance is often a bottleneck in various real-world applications. Therefore, we address this by designing an efficient subgraph isomorphism algorithm leveraging features of GPU architecture, such as massive parallelism and memory hierarchy. Existing GPU-based solutions adopt a two-step output scheme, performing the same join process twice in order to write intermediate results concurrently. They also lack GPU architecture-aware optimizations that allow scaling to large graphs. In this paper, we propose a GPU-friendly subgraph isomorphism algorithm, GSI. Different from existing edge join-based GPU solutions, we propose a Prealloc-Combine strategy based on the vertex-oriented framework, which avoids joining-twice in existing solutions. Also, a GPU-friendly data structure (called PCSR) is proposed to represent an edge-labeled graph. Extensive experiments on both synthetic and real graphs show that GSI outperforms the state-of-the-art algorithms by up to several orders of magnitude and has good scalability with graph size scaling to hundreds of millions of edges.

cs.DB

An Accurate and Transferable Machine-Learning Interatomic Potential for Silicon

The development of modern ab initio methods has rapidly increased our understanding of physics, chemistry and materials science. Unfortunately, intensive ab initio calculations are intractable for large and complex systems. On the other hand, empirical force fields are less accurate with poor transferability even though they are efficient to handle large and complex systems. The recent development of machine-learning based neural-network (NN) for local atomic environment representation of density functional theory (DFT) has offered a promising solution to this long-standing challenge. Si is one of the most important elements in science and technology, however, an accurate and transferable interatomic potential for Si is still lacking. Here, we develop a generalized NN potential for Si, which correctly predicts the Si(111)-(7x7) ground-state surface reconstruction for the first time and accurately reproduces the DFT results in a wide range of complex Si structures. We envision similar developments will be made for a wide range of materials systems in the near future.

cond-mat.mtrl-sci

Ubiquitous Ideal Spin-Orbit Coupling in a Screw Dislocation in Semiconductors

We theoretically demonstrate that screw dislocation (SD), a 1D topological defect widely present in semiconductors, exhibits ubiquitously a new form of spin-orbit coupling (SOC) effect. Differing from the widely known conventional 2D Rashba-Dresselhaus (RD) SOC effect that typically exists at surfaces/interfaces, the deep-level nature of SD-SOC states in semiconductors readily makes it an ideal SOC. Remarkably, the spin texture of 1D SD-SOC, pertaining to the inherent symmetry of SD, exhibits a significantly higher degree of spin coherency than the 2D RD-SOC. Moreover, the 1D SD-SOC can be tuned by ionicity in compound semiconductors to ideally suppress spin relaxation, as demonstrated by comparative first-principles calculations of SDs in Si/Ge, GaAs, and SiC. Our findings therefore open a new door to manipulating spin transport in semiconductors by taking advantage of an otherwise detrimental topological defect.

cond-mat.mtrl-sci

On the connection problem for nonlinear differential equation

We consider the connection problem of the second nonlinear differential equation \begin{equation} \label{eq:1} \Phi''(x)=(\Phi'^2(x)-1)\cot\Phi(x)+ \frac{1}{x}(1-\Phi'(x)) \end{equation} subject to the boundary condition $\Phi(x)=x-ax^2+O(x^3)$ ($a\geq0$) as $x\to0$. In view of that equation (1) is equivalent to the fifth Painlev\'e (PV) equation after a M\"obius transformation, we are able to study the connection problem of equation (1) by investigating the corresponding connection problem of PV. Our research technique is based on the method of uniform asymptotics presented by Bassom el at. The monotonically solution on real axis of equation (1) is obtained, the explicit relation (connection formula) between the constants in the solution and the real number $a$ is also obtained. This connection formulas have been established earlier by Suleimanov via the isomonodromy deformation theory and the WKB method, and recently are applied for studying level spacing functions.

math.CA