arXiv ScienceSearch

arXiv subjects

Ting Guo

Publications and source records attributed to Ting Guo.

15 recordsLinked to original sources

$d$-spacing distributions as a probe of nematoelastic response in iron-based superconductors

Electronic nematicity in iron-based superconductors (FeSCs) couples bilinearly to orthorhombic strain, allowing nematic correlations to appear in the lattice response. Here we use neutron Larmor diffraction to measure the temperature-dependent distribution of relative $d$ spacings in electron-doped Ba(Fe$_{1-x}$Co$_x$)$_2$As$_2$, hole-doped Ba$_{0.83}$K$_{0.17}$Fe$_2$As$_2$, FeSe, and Fe$_{1.07}$Te. In Ba(Fe$_{1-x}$Co$_x$)$_2$As$_2$ crystals without intentionally applied uniaxial stress, the in-plane distribution width, $\varepsilon_{\rm FWHM}$, increases on cooling in the tetragonal phase and can be described phenomenologically by a Curie--Weiss-like form. The fitted scale $T^*$ decreases with Co doping and evolves similarly to the nematic phase diagram inferred from elastoresistance, although the two experiments probe different response functions. Related broadening in Ba$_{0.83}$K$_{0.17}$Fe$_2$As$_2$ and FeSe supports extending this interpretation beyond electron-doped BaFe$_2$As$_2$. By contrast, Fe$_{1.07}$Te shows no extended Curie--Weiss-like regime without applied stress, whereas uniaxial pressure produces a strongly anisotropic broadening that can contain contributions from both the field-biased lattice response and inhomogeneous loading. A mean-field model with bilinear nematoelastic coupling and spatially varying symmetry-breaking stress explains the Curie--Weiss-like broadening in terms of the renormalized orthorhombic compliance. Neutron Larmor diffraction therefore provides a bulk-sensitive probe of nematic-related lattice broadening that complements electronic and elastic measurements.

cond-mat.supr-con

Every Activation Boosted: Scaling General Reasoner to 1 Trillion Open Language Foundation

We introduce Ling 2.0, a series reasoning-oriented language foundation built upon the principle that every activation boosts reasoning capability. Designed to scale from tens of billions to one trillion parameters under a unified Mixture-of-Experts (MoE) paradigm, Ling 2.0 emphasizes high sparsity, cross-scale consistency, and efficiency guided by empirical scaling laws. The series includes three non-thinking (instruct) models - Ling-mini-2.0, Ling-flash-2.0, and Ling-1T - ranging from 16B to 1T total parameters and achieving up to 7-fold active-compute efficiency compared with dense counterparts. Ling 2.0 integrates coordinated innovations across model architecture, pre-training, post-training, and infrastructure: a high-sparsity MoE with MTP for efficient reasoning, reasoning-oriented data and mid-training CoT activation, reinforcement-based fine-tuning (DFT, Evo-CoT), and full-scale FP8 training with fine-grained heterogeneous pipelines. At the trillion scale, Ling-1T establishes a new Pareto frontier of reasoning accuracy versus computational efficiency, demonstrating that sparse activation, when properly aligned with reasoning objectives, enables scalable and efficient intelligence. Collectively, Ling 2.0 provides a coherent, open, and efficient foundation for advancing future reasoning and thinking models, including the Ring series built upon the same base.

cs.CL

Dual role of stripe phase on superconducting correlation in a bilayer square lattice

While the stripe phase has been observed not only in monolayer cuprates but also in bilayer cuprates, research on its behavior in bilayer cuprates has been limited. Using constrained path quantum Monte Carlo, we explore the effect of stripes on the bilayer square lattice. We find the system exhibits short-range antiferromagnetism, which is enhanced by stripes and is strongest when the electron density of the interstriped rows reaches half-filling. The hole doping concentration plays a crucial role in the interaction between stripes and superconductivity. The $d$-wave pairing is enhanced by stripe potential $V_0$ at the hole doping $\delta_h=1/4$, whereas it is suppressed by stripe potential $V_0$ at the hole doping $\delta_h=1/8$. We elucidate this phenomenon through an analysis of the magnetism of the interstriped rows. Furthermore, the effective $d$-wave pairing is stronger in the bilayer model compared to the monolayer model when stripes are introduced on the square lattice. Overall, our unbiased numerical simulations provide a further understanding of the crossed bilayer square lattice model.

cond-mat.str-el

Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM

Recent advancements in code large language models (LLMs) have demonstrated remarkable capabilities in code generation and understanding. It is still challenging to build a code LLM with comprehensive performance yet ultimate efficiency. Many attempts have been released in the open source community to break the trade-off between performance and efficiency, such as the Qwen Coder series and the DeepSeek Coder series. This paper introduces yet another attempt in this area, namely Ling-Coder-Lite. We leverage the efficient Mixture-of-Experts (MoE) architecture along with a set of high-quality data curation methods (especially those based on program analytics) to build an efficient yet powerful code LLM. Ling-Coder-Lite exhibits on-par performance on 12 representative coding benchmarks compared to state-of-the-art models of similar size, such as Qwen2.5-Coder-7B and DeepSeek-Coder-V2-Lite, while offering competitive latency and throughput. In practice, we achieve a 50\% reduction in deployment resources compared to the similar-sized dense model without performance loss. To facilitate further research and development in this area, we open-source our models as well as a substantial portion of high-quality data for the annealing and post-training stages. The models and data can be accessed at~\url{https://huggingface.co/inclusionAI/Ling-Coder-lite}.

cs.LG

Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs

In this technical report, we tackle the challenges of training large-scale Mixture of Experts (MoE) models, focusing on overcoming cost inefficiency and resource limitations prevalent in such systems. To address these issues, we present two differently sized MoE large language models (LLMs), namely Ling-Lite and Ling-Plus (referred to as "Bailing" in Chinese, spelled B\v{a}il\'ing in Pinyin). Ling-Lite contains 16.8 billion parameters with 2.75 billion activated parameters, while Ling-Plus boasts 290 billion parameters with 28.8 billion activated parameters. Both models exhibit comparable performance to leading industry benchmarks. This report offers actionable insights to improve the efficiency and accessibility of AI development in resource-constrained settings, promoting more scalable and sustainable technologies. Specifically, to reduce training costs for large-scale MoE models, we propose innovative methods for (1) optimization of model architecture and training processes, (2) refinement of training anomaly handling, and (3) enhancement of model evaluation efficiency. Additionally, leveraging high-quality data generated from knowledge graphs, our models demonstrate superior capabilities in tool use compared to other models. Ultimately, our experimental findings demonstrate that a 300B MoE LLM can be effectively trained on lower-performance devices while achieving comparable performance to models of a similar scale, including dense and MoE models. Compared to high-performance devices, utilizing a lower-specification hardware system during the pre-training phase demonstrates significant cost savings, reducing computing costs by approximately 20%. The models can be accessed at https://huggingface.co/inclusionAI.

cs.LG

Context-Aware Lifelong Sequential Modeling for Online Click-Through Rate Prediction

Lifelong sequential modeling (LSM) is becoming increasingly critical in social media recommendation systems for predicting the click-through rate (CTR) of items presented to users. Central to this process is the attention mechanism, which extracts interest representations with respect to candidate items from the user sequence. Typically, attention mechanisms operate in a point-wise manner, focusing solely on the relevance of individual items in the sequence to the candidate item. In contrast, context-aware LSM aims to also consider adjacent items in the user behavior sequence to better assess the importance of each item. In this paper, we propose the Context-Aware Interest Network (CAIN), which utilizes the Temporal Convolutional Network (TCN) to create context-aware representations for each item throughout the lifelong sequence. These enhanced representations are then used in the attention mechanism instead of the original item representations to derive context-aware interest representations. Building upon this TCN framework, we propose the Multi-Scope Interest Aggregator (MSIA) module, which incorporates multiple TCN layers and their corresponding attention modules to capture interest representations across varying context scopes. Furthermore, we introduce the Personalized Extractor Generation (PEG) module, which generates convolution filters based on users' basic profile features. These personalized filters are then used in the TCN layers instead of the original global filters to generate more user-specific representations. We conducted extensive experiments on both a public dataset and an industrial dataset from the WeChat Channels platform. The results demonstrate that CAIN outperforms existing methods in terms of prediction accuracy and online performance metrics.

cs.IR

Hecke growth diagrams, and maximal increasing and decreasing sequences in fillings of stack polyominoes

We establish a bijection between $01$-fillings of stack polyominoes with at most one $1$ per column and labelings of the corners along the top-right border of stack polyominoes. These labellings indicate the lengths of the longest increasing and decreasing chains of the largest rectangular region below and to the left of the corners. Our results provide an alternative proof of Guo and Poznanovi\'c's theorem on the lengths of the longest increasing and decreasing chains have a symmetric joint distribution over $01$-fillings of stack polyomino. Moreover, our results offer new perspective to Chen, Guo and Pang's result on the crossing number and the nesting number have a symmetric joint distribution over linked partitions. In particular, our construction generalizes the growth diagram techniques of Rubey for the $01$-fillings of stack polyominoes with at most one $1$ per column and row.

math.CO

Existence and qualitative properties of solutions for a Choquard-type equation with Hardy potential

In this paper, we study the existence and qualitative properties of positive solutions to a Choquard-type equation with Hardy potential. We develop a nonlocal version of concentration-compactness principle involving the Hardy potential to study the existence and the asymptotic behavior of positive solutions by transforming the original problem into a new nonlocal problem in the weighted Sobolev space. Moreover, we obtain the symmetry of solutions by using the moving plane method.

math.AP

Different phase leads to different transport behavior in Pb$_9$Cu(PO$_4$)$_6$O compounds

The recent claimed room-temperature superconductivity in Cu-doped lead apatite at ambient pressure are under highly debate. To identify its physical origin, we studied the crystal structures, energy band structures, lattice dynamics and magnetic properties of the parent Pb$_{10}$(PO$_4$)$_6$O compound, in which two different phases of the LK-99 compound are analyzed in detail. Our results show that the Pb$_{10}$(PO$_4$)$_6$O compound is an indirect band gap semiconductor, where Cu doping at the 4$f$ site of Pb leads to a semiconducting to half-metallic transition. Two half-filled flat bands spanning the Fermi energy levels are present in the 4$f$-phase of LK-99, which are mainly formed by hybridization of the $d_{x^2-y^2}$ and $d_{zy}$ orbitals of Cu with the 2$p$ orbitals of O. In addition, 6$h$-phase of LK-99 always has spin polarity at the bottom of the conduction band and at the top of the valence band, making the material a bipolar magnetic semiconductor. Our results are basically consistent with the recent experimental transport properties of LK-99 posted on arXiv:2308.05778.

cond-mat.supr-con

Spatio-Temporal Contrastive Learning Enhanced GNNs for Session-based Recommendation

Session-based recommendation (SBR) systems aim to utilize the user's short-term behavior sequence to predict the next item without the detailed user profile. Most recent works try to model the user preference by treating the sessions as between-item transition graphs and utilize various graph neural networks (GNNs) to encode the representations of pair-wise relations among items and their neighbors. Some of the existing GNN-based models mainly focus on aggregating information from the view of spatial graph structure, which ignores the temporal relations within neighbors of an item during message passing and the information loss results in a sub-optimal problem. Other works embrace this challenge by incorporating additional temporal information but lack sufficient interaction between the spatial and temporal patterns. To address this issue, inspired by the uniformity and alignment properties of contrastive learning techniques, we propose a novel framework called Session-based Recommendation with Spatio-Temporal Contrastive Learning Enhanced GNNs (RESTC). The idea is to supplement the GNN-based main supervised recommendation task with the temporal representation via an auxiliary cross-view contrastive learning mechanism. Furthermore, a novel global collaborative filtering graph (CFG) embedding is leveraged to enhance the spatial view in the main task. Extensive experiments demonstrate the significant performance of RESTC compared with the state-of-the-art baselines e.g., with an improvement as much as 27.08% gain on HR@20 and 20.10% gain on MRR@20.

cs.IR

Quantum Monte Carlo study of superconductivity in rhombohedral trilayer graphene under an electric field

By using the constrained-phase quantum Monte Carlo method, we performed a systematic study of the ground state of the half filled Hubbard model for a trilayer honeycomb lattice. We analyze the effect of the perpendicular electric field on the electronic structure, magnetic property and pairing correlations. It is found that the antiferromagnetism is suppressed by the perpendicular electric field, especially the long-range parts, and the dominant magnetic fluctuations are still antiferromagnetic. The electronic correlation drives a $d+id$ superconducting pairing to be dominant over other pairing patterns among various electric fields and interaction strengths. We also found that the $d+id$ pairing correlation is greatly enhanced as the on-site Coulomb interaction is increased. Our intensive numerical results may unveil the nature of the recently observed superconductivity in rhombohedral trilayer graphene under an electric field.

cond-mat.str-el

Enhancement of $d$-wave pairing in the striped phase with the nearest neighbour attraction

Recently, the experimental results by the angle-resolved photoemission spectroscopy suggested that an additional strong nearest neighbor attraction in the Hubbard model might be significant to describe the properties of doped cuprates more accurately. The stripe-ordered patterns, formed by the inhomogeneous distribution of spin, charge and pairing correlations in the CuO$_{2}$ planes, is a known feature of doped cuprates. In this work, the effect of the nearest neighbor attraction and the stripe phase are examined by using the constrained path quantum Monte Carlo method within the repulsive Hubbard model on two-dimensional square lattice. The ground state spin correlations along and cross the stripe regions, and the $d$-wave pairing correlation are calculated. It is found that the spin-spin correlation is the highest when the interstripe region is fairly close to half-filling, and $d$-wave superconducting correlation on neighboring sites could be enhanced in the presence of stripe pattern and strong nearest neighbor attraction, which reveals their crucial roles on superconductivity in the doped cuprates.

cond-mat.str-el

Hecke insertion and maximal increasing and decreasing sequences in fillings of stack polyominoes

We prove that the number of 01-fillings of a given stack polyomino (a polyomino with justified rows whose lengths form a unimodal sequence) with at most one 1 per column which do not contain a fixed-size northeast chain and a fixed-size southeast chain, depends only on the set of row lengths of the polyomino. The proof is via a bijection between fillings of stack polyominoes which differ only in the position of one row and uses the Hecke insertion algorithm by Buch, Kresch, Shimozono, Tamvakis, and Yong and the jeu de taquin for increasing tableaux of Thomas and Yong. Moreover, our bijection gives another proof of the result by Chen, Guo, and Pang that the crossing number and the nesting number have a symmetric joint distribution over linked partitions.

math.CO

Superconvergence of differential structure for finite element methods on perturbed surface meshes

Superconvergence of differential structure on discretized surfaces is studied in this paper. The newly introduced geometric supercloseness provides us with a fundamental tool to prove the superconvergence of gradient recovery on deviated surfaces. An algorithmic framework for gradient recovery without exact geometric information is introduced. Several numerical examples are documented to validate the theoretical results.

math.NA

On (shape-)Wilf-equivalence for words

Stankova and West showed that for any non-negative integer $s$ and any permutation $\gamma$ of $\{4,5,\dots,s+3\}$ there are as many permutations that avoid $231\gamma$ as there are that avoid $312\gamma$. We extend this result to the setting of words.

math.CO