arXiv ScienceSearch

arXiv subjects

Xinyang Liu

Publications and source records attributed to Xinyang Liu.

At least 19 recordsLinked to original sources

Do Uncertainty Signals Help? A Systematic Study of Uncertainty-Aware Decoding with Rollback Mechanisms

Prediction uncertainty is a widely adopted metric for quantifying model confidence, with downstream applications spanning model explanation, data selection, and prediction rollback. Despite its demonstrated utility, the potential of uncertainty quantification to enhance code generation in large language models (LLMs) remains largely underexplored, raising a critical question: to what extent can uncertainty serve as an effective signal for improving LLM-based code generation? To answer this question, we study uncertainty-aware rollback decoding, an inference-time strategy that uses uncertainty signals to identify unreliable generation regions and roll back to earlier valid prefixes without retraining the model. We evaluate this framework on seven code LLMs, five code generation benchmarks, and eight token-level uncertainty signals under a unified decoding setup. Our results show that the complete rollback framework improves over equal-budget restart across the evaluated benchmarks and model settings, with gains of up to 0.26 in pass@1 and 0.35 in AvgTestPassRate on functional code generation benchmarks, and an absolute improvement of up to 6.4\% in Patch-Aligned Safe Rate on Dsec-Python. Among the evaluated signals, information-theoretic measures such as token entropy and negative log-likelihood show the most favorable overall trend, frequently achieving the best or near-best results on standard benchmarks. A component-controlled ablation further shows that feedback-guided rollback provides the main improvement, while uncertainty localization provides an additional gain when checking, budget, rollback, and branch decay are held fixed.

cs.LG

Overlooked weak structural connections support human cognition under nonlinear connectome scaling

Human cognition depends on large scale communication constrained by white matter architecture. Although weak connections are abundant in mammalian connectomes, they have long been treated as noise and downweighted because of tractography uncertainty in the human brain, and their relevance to human cognition and large scale functional organization remains unresolved. Across multiple datasets and tractography pipelines, we show that, when tractography derived connectivity weights are interpreted through a nonlinear weighting framework, weak connections make measurable contributions to cognitive prediction, functional connectivity simulation, and structure-function coupling. These effects are selective: nonlinear weighting improves the prediction of general cognitive ability and memory more than that of crystallized intelligence or processing speed, consistent with the notion that weak connections preferentially expand the modal repertoire of brain networks to enhance both large scale integration and fine grained segregation, thereby supporting the functional balance essential for diverse cognitive abilities. Importantly, these effects are replicated in a reliability aware connectome generated by integrating two post tractography filtering methods, in which preserving weak links consistently outperforms conventional thresholding strategies. Finally, we show that weak connections contain functionally informative subsets organized along systems level and transcriptomic gradients. In particular, a specific class of weak connections, predominantly linking visual and motor systems with limbic regions and characterized by negative gene coexpression, exerts a disproportionately large influence on brain function.

q-bio.NC

Ising Supercriticality and Universal Magnetocalorics in Spiral Antiferromagnet Nd$_3$BWO$_9$

The celebrated analogy between the pressure-temperature phase diagram of a liquid-gas system and the field-temperature phase diagram of a ferromagnet has long been a cornerstone for understanding universality of phase transitions and critical phenomena. Here we extend this analogy to a highly frustrated antiferromagnet, the spiral Ising compound Nd$_3$BWO$_9$ with kagome layers. In its phase diagram, we identify a metamagnetic transition line with a critical endpoint (CEP) located at $μ_0H_{\mathrm{c}} \simeq 1.04$ T and $T_{\mathrm{c}} \simeq 0.3$ K. Above the CEP, an Ising supercritical regime emerges with crossover lines that follow a universal scaling law, as evidenced by the specific heat, magnetic susceptibility, and magnetocaloric measurements. Remarkably, we observe highly sensitive magnetic cooling near the emergent CEP, characterized by a divergent magnetic Grüneisen ratio $Γ_H \propto 1/t^{β+γ-1}$, with $β+ γ\simeq 1.563$ the sum of critical exponents of the 3D Ising universality class and $t \equiv (T-T_{\rm c})/T_{\rm c}$ the reduced temperature. Adiabatic demagnetization from 2 K and 4 T reaches a minimum temperature of 195 mK, via a self-cascading process that combines supercritical and topological cooling. Our findings open a new avenue for studying supercritical phenomena and magnetic refrigeration with the frustrated rare-earth compounds RE$_3$BWO$_9$ and, more broadly, in Ising-anisotropic antiferromagnets such as spin ices.

cond-mat.str-el

FFR: Forward-Forward Learning for Regression

The Forward-Forward (FF) algorithm offers a computationally efficient and biologically plausible alternative to backpropagation (BP) by training neural networks through purely local, layer-wise optimization. However, FF is inherently designed for classification via contrastive positive-negative sample pairs, and extending it to regression poses fundamental challenges: continuous target space lack natural "opposites" for contrastive learning, and the standard goodness function carries no information about target magnitude or ordering. We propose FFR (Forward-Forward for Regression), to our knowledge, the first framework to extend FF to real-world regression and demonstrate competitive performance across diverse real-world datasets. FFR introduces three key innovations: (1) an ordinal competitive goodness function that replaces contrastive pairs with competitive learning between partitioned neuron groups under distance-aware ordinal supervision; (2) a stratified ladder architecture where shallow layers learn coarse ordinal discrimination and deeper layers refine into fine-grained regression, with multi-scale feature aggregation for inter-layer collaboration; and (3) hierarchical prediction with uncertainty estimation, where multi-scale predictors jointly provide robust predictions and prediction confidence as a free-lunch. Extensive experimental results show FFR recovers on average 98.6% of BP's accuracy across five real-world regression benchmarks while reducing peak training memory to only 27% of BP's at depth 8 and 8% at depth 32, with per-iteration time around 72% of BP's, and substantially outperforms all BP-free competitors.

cs.LG

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery

Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-The-Art (SOTA) performance, creating a growing need for systems that can accelerate the full pipeline of empirical model optimization. In this work, we introduce AutoSOTA, an end-to-end automated research system that advances the latest SOTA models published in top-tier AI papers to reproducible and empirically improved new SOTA models. We formulate this problem through three tightly coupled stages: resource preparation and goal setting; experiment evaluation; and reflection and ideation. To tackle this problem, AutoSOTA adopts a multi-agent architecture with eight specialized agents that collaboratively ground papers to code and dependencies, initialize and repair execution environments, track long-horizon experiments, generate and schedule optimization ideas, and supervise validity to avoid spurious gains. We evaluate AutoSOTA on recent research papers collected from eight top-tier AI conferences under filters for code availability and execution cost. Across these papers, AutoSOTA achieves strong end-to-end performance in both automated replication and subsequent optimization. Specifically, it successfully discovers 105 new SOTA models that surpass the original reported methods, averaging approximately five hours per paper. Case studies spanning LLM, NLP, computer vision, time series, and optimization further show that the system can move beyond routine hyperparameter tuning to identify architectural innovation, algorithmic redesigns, and workflow-level improvements. These results suggest that end-to-end research automation can serve not only as a performance optimizer, but also as a new form of research infrastructure that reduces repetitive experimental burden and helps redirect human attention toward higher-level scientific creativity.

cs.CL

Noise-like pulse laser source with ultrabroadband tunability and coherence-limited sub-structure

High brightness and low coherence laser sources with wideband tunability are essential for many full-field imaging applications aiming for high contrast and speckle free performance. However, this combination of parameters is challenging to achieve. The current solutions focus on decreasing spatial coherence or generation of time-varying speckle patterns, while suppression of temporal coherence typically compromises brightness. Here we demonstrate a wideband pulsed laser source with low temporal coherence and the absence of phase correlation between pulses as an alternative approach with simultaneous time and frequency diversity. The full gain spectrum of a Tm doped fiber laser (1650 nm 2000 nm) is operated in a tunable noise like pulse regime, which by nature is composed of countless structured elementary events with uncorrelated phases randomly varying from bunch to bunch. The measured spectral widths range from 13.8 nm to 18.8 nm, while the average output power varies between 63.3 mW and 213 mW. Numerical simulations reveal that temporal coherence decreases significantly with increasing optical gain, dropping from near unity at low gain to approximately 0.2 at high gain. The startup dynamics of the noise like pulse laser are experimentally studied using the dispersive Fourier transformation (DFT) method. Based on single shot spectra and frequency resolved optical gating traces, the coherence properties of the laser are further analyzed by calculating the mutual coherence function and cross-spectral density. The noise like pulse laser exhibits a coherence time of approximately 100 fs and an average pulse burst duration of about 40 ps in the high-gain regime.

physics.optics

Disentangled Generative Graph Representation Learning

Recently, generative graph models have shown promising results in learning graph representations through self-supervised methods. However, most existing generative graph representation learning (GRL) approaches rely on random masking across the entire graph, which overlooks the entanglement of learned representations. This oversight results in non-robustness and a lack of explainability. Furthermore, disentangling the learned representations remains a significant challenge and has not been sufficiently explored in GRL research. Based on these insights, this paper introduces DiGGR (Disentangled Generative Graph Representation Learning), a self-supervised learning framework. DiGGR aims to learn latent disentangled factors and utilizes them to guide graph mask modeling, thereby enhancing the disentanglement of learned representations and enabling end-to-end joint learning. Extensive experiments on 11 public datasets for two different graph learning tasks demonstrate that DiGGR consistently outperforms many previous self-supervised methods, verifying the effectiveness of the proposed approach.

cs.LG

Damping dynamics of the centroid oscillation of a relativistic laser pulse in a plasma channel

The centroid oscillation of an offset laser pulse propagating in a preformed plasma channel is investigated through theoretical analysis and three-dimensional particle-in-cell simulations. For non-relativistic laser pulses, the mode leakage of a finite channel and the temporal walk-off between the fundamental and high order modes of a finite-duration laser induce a decay in the laser centroid oscillation. An analytical model characterizing these decay mechanisms is derived and validated by simulations. For relativistic laser pulses, the slice-based centroid oscillation frequency develops an axial chirp due to relativistic channel modification and photon deceleration. This chirp leads to phase mixing across different axial slices of the pulse, resulting in a rapid damping of the overall centroid oscillation. Understanding this oscillation damping is crucial for mitigating electron beam pointing jitter and maintaining beam quality in high-energy, channel-guided laser wakefield accelerators.

physics.acc-ph

Route Experts by Sequence, not by Token

Mixture-of-Experts (MoE) architectures scale large language models (LLMs) by activating only a subset of experts per token, but the standard TopK routing assigns the same fixed number of experts to all tokens, ignoring their varying complexity. Prior adaptive routing methods introduce additional modules and hyperparameters, often requiring costly retraining from scratch. We propose Sequence-level TopK (SeqTopK), a minimal modification that shifts the expert budget from the token level to the sequence level. By selecting the top $T \cdot K$ experts across all $T$ tokens, SeqTopK enables end-to-end learned dynamic allocation -- assigning more experts to difficult tokens and fewer to easy ones -- while preserving the same overall budget. SeqTopK requires only a few lines of code, adds less than 1% overhead, and remains fully compatible with pretrained MoE models. Experiments across math, coding, law, and writing show consistent improvements over TopK and prior parameter-free adaptive methods, with gains that become substantially larger under higher sparsity (up to 16.9%). These results highlight SeqTopK as a simple, efficient, and scalable routing strategy, particularly well-suited for the extreme sparsity regimes of next-generation LLMs. Code is available at https://github.com/Y-Research-SBU/SeqTopK.

cs.LG

Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct

Fast and high-quality language generation is the holy grail that people pursue in the age of AI. In this work, we introduce Discrete Diffusion Divergence Instruct (DiDi-Instruct), a training-based method that initializes from a pre-trained diffusion large language model (dLLM) and distills a few-step student for fast generation. The model distilled with DiDi-Instruct matches or surpasses its dLLM teacher and the GPT-2 baseline while providing up to 64$\times$ acceleration. The theoretical foundation of DiDi-Instruct is a novel framework based on integral KL-divergence minimization, which leads to a practical training algorithm. We further introduce grouped reward normalization, intermediate-state matching, and the reward-guided ancestral sampler to improve training stability, model coverage, and inference quality. On the OpenWebText benchmark, DiDi-Instruct achieves perplexity ranging from 62.2 (8 NFEs) to 18.4 (128 NFEs), outperforming prior accelerated dLLMs and the GPT-2 baseline. These gains incur a negligible entropy loss (around $1$%) and reduce additional training wall-clock time by more than $20\times$ compared to competing dLLM distillation methods. We further validate the robustness and effectiveness of DiDi-Instruct through extensive ablation studies, model scaling, downstream task evaluations, and unconditional protein sequence generation. In conclusion, DiDi-Instruct enables efficient and effective distillation for language generation in the blink of an eye.

cs.CL

Giant Magnetocaloric Effect in a High-Spin Shastry-Sutherland Dipolar Magnet

The Shastry-Sutherland lattice is a prototypical frustrated quantum magnet. It is notable for its exactly solvable dimer-singlet ground state and hosts a wealth of magnetic phenomena under external fields. Here, this work investigates the high-spin (S = 7/2) Eu-based magnet Eu2MgSi2O7 (EMSO) using low-temperature magnetothermal measurements and Monte Carlo simulations, revealing a giant magnetocaloric effect (MCE) in this Shastry-Sutherland compound. The entropy change peak value is found to be 55.0 J kg-1 K-1 under a field change of B = 0-4 T, approximately 1.5 times larger than the commercial Gd3Ga5O12 (GGG). Adiabatic demagnetization refrigeration achieves a lowest temperature of 151 mK, deeply into the sub-Kelvin regime. Furthermore, a distinctive cooling effect persists below about 1 T, a characteristic absent for conventional magnetic coolants. A dipolar Shastry-Sutherland model is introduced as a minimal model to describe this system; in particular, the experimentally revealed 1/3 magnetization pseudo-plateau can be ascribed to the presence of dipolar couplings between Eu2+ ions, further stabilized by the thermal fluctuations, explaining the persistent cooling effect. This work establishes EMSO as a novel platform for exploring the dipolar Shastry-Sutherland system and for sub-Kelvin adiabatic demagnetization refrigeration.

cond-mat.mtrl-sci

GDEPO: Group Dual-dynamic and Equal-right Advantage Policy Optimization with Enhanced Training Data Utilization for Sample-Constrained Reinforcement Learning

Automated Theorem Proving (ATP) represents a fundamental challenge in Artificial Intelligence (AI), requiring the construction of machine-verifiable proofs in formal languages such as Lean to evaluate AI reasoning capabilities. Reinforcement learning (RL), particularly the high-performance Group Relative Policy Optimization (GRPO) algorithm, has emerged as a mainstream approach for this task. However, in ATP scenarios, GRPO faces two critical issues: when composite rewards are used, its relative advantage estimation may conflict with the binary feedback from the formal verifier; meanwhile, its static sampling strategy may discard entire batches of data if no valid proof is found, resulting in zero contribution to model updates and significant data waste. To address these limitations, we propose Group Dual-dynamic and Equal-right-advantage Policy Optimization (GDEPO), a method incorporating three core mechanisms: 1) dynamic additional sampling, which resamples invalid batches until a valid proof is discovered; 2) equal-right advantage, decoupling the sign of the advantage function (based on correctness) from its magnitude (modulated by auxiliary rewards) to ensure stable and correct policy updates; and 3) dynamic additional iterations, applying extra gradient steps to initially failed but eventually successful samples to accelerate learning on challenging cases. Experiments conducted on three datasets of varying difficulty (MinF2F-test, MathOlympiadBench, PutnamBench) confirm the effectiveness of GDEPO, while ablation studies validate the necessity of its synergistic components. The proposed method enhances data utilization and optimization efficiency, offering a novel training paradigm for ATP.

cs.AI

Canted ferromagnetic order in a distorted triangular-lattice magnet Na$_2$SrCo(VO$_4$)$_2$

Triangular-lattice cobaltates with glaserite-type $X_2Y$Co($T$O$_4)_2$ structure provide an ideal platform to investigate intriguing quantum magnetism. Here we report a comprehensive study of the structural and magnetic properties of a triangular-lattice cobalt vanadate $\rm Na_2SrCo(VO_4)_2$. Room-temperature x-ray and neutron powder diffraction confirm that $\rm Na_2SrCo(VO_4)_2$ crystallizes in the monoclinic $P2_1/c$ space group with slightly distorted triangular layers of $\rm Co^{2+}$ ions. Magnetization measurements reveal a ferromagnetic transition at $T\rm_C \approx 3.4~{\rm K}$, where a sharp $λ$-type anomaly is observed in the specific heat. The magnetic entropy recovered up to 55 K approaches 90$\%$ of $R{\rm ln}2$, supporting an effective spin-1/2 state of Co$^{2+}$ ions at low temperature. Neutron diffraction at 2.3 K (below $T_{\rm C}$) further confirms a long-range canted ferromagnetic order with the Co$^{2+}$ moments aligned in the $ac$ plane and the ordered moment size of $\sim$ 2.6 $μ\rm_{B}$. Comparing with its sister compounds with a trigonal symmetry, $\rm Na_2BaCo(VO_4)_2$ with a collinear ferromagnetic structure and the recently discovered spin supersolid candidate $\rm Na_2BaCo(PO_4)_2$ with a distinct Y-like antiferromagnetic ground state, this study indicates the decisive role of the $T{\rm O_4}$ tetrahedra in tuning exchange interactions and contrasting magnetic behaviors of these glaserite-structure compounds.

cond-mat.str-el

OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists

With the rapid development of Large Language Models (LLMs), AI agents have demonstrated increasing proficiency in scientific tasks, ranging from hypothesis generation and experimental design to manuscript writing. Such agent systems are commonly referred to as "AI Scientists." However, existing AI Scientists predominantly formulate scientific discovery as a standalone search or optimization problem, overlooking the fact that scientific research is inherently a social and collaborative endeavor. Real-world science relies on a complex scientific infrastructure composed of collaborative mechanisms, contribution attribution, peer review, and structured scientific knowledge networks. Due to the lack of modeling for these critical dimensions, current systems struggle to establish a genuine research ecosystem or interact deeply with the human scientific community. To bridge this gap, we introduce OmniScientist, a framework that explicitly encodes the underlying mechanisms of human research into the AI scientific workflow. OmniScientist not only achieves end-to-end automation across data foundation, literature review, research ideation, experiment automation, scientific writing, and peer review, but also provides comprehensive infrastructural support by simulating the human scientific system, comprising: (1) a structured knowledge system built upon citation networks and conceptual correlations; (2) a collaborative research protocol (OSP), which enables seamless multi-agent collaboration and human researcher participation; and (3) an open evaluation platform (ScienceArena) based on blind pairwise user voting and Elo rankings. This infrastructure empowers agents to not only comprehend and leverage human knowledge systems but also to collaborate and co-evolve, fostering a sustainable and scalable innovation ecosystem.

cs.CY

Route-and-Reason: Scaling Large Language Model Reasoning with Reinforced Model Router

Chain-of-thought has been proven essential for enhancing the complex reasoning abilities of Large Language Models (LLMs), but it also leads to high computational costs. Recent advances have explored the method to route queries among multiple models and proved it as a promising approach. However, previous works directly operate at the task level, i.e., assigning user queries to suitable LLMs, which does not allow hybrid LLMs to truly collaborate on finer-grained sub-tasks. Collaboration at the level of intermediate reasoning steps (thoughts) could enable more efficient coordination, but it also poses significant challenges for router scheduling, placing immense demands on the quality of task decomposition and the precision of the router. To address this, we propose R2-Reasoner, a novel framework centered around a Reinforced Model Router designed to efficiently scale LLM reasoning. This router orchestrates collaboration across nine heterogeneous models, whose parameter scales range from less than 1B to hundreds of billions, by first breaking down a complex query into subtasks with a decomposer, and then assigning each subtask to the optimal model with a subtask allocator, balancing performance with cost. Training this router involves a two-stage alternating process for the decomposer and the allocator, integrating supervised fine-tuning with reinforcement learning to enable effective self-supervised refinement. Extensive experiments across six challenging reasoning benchmarks demonstrate that R2-Reasoner reduces API costs by 84.46% compared with state-of-the-art baselines while maintaining competitive reasoning accuracy. Our framework paves the way for the development of more scalable and efficient reasoning systems. Our code is open-source at https://anonymous.4open.science/r/R2_Reasoner.

cs.CL

Quantum fluctuations associated with first-order magnetic transition in a frustrated kagome lattice antiferromagnet

Intense quantum fluctuations arising from geometrical frustrations in kagome-lattice magnets provide a feasible approach to exotic quantum states. Here, we document an unexpected isosymmetric first-order magnetic transition in the recently synthesized frustrated kagome-lattice antiferromagnet Nd3ScBi5, which is characterized by significant latent heat and a pronounced magnetocaloric effect, as well as discontinuous Raman shifts and negligible hysteresis. Employing the magnetocaloric effect as a detection method, in conjunction with systematical field-dependent physical properties, we uncover a distinctive 1/2 magnetization plateau phase with significant quantum fluctuations. Our study unveils Nd3ScBi5 as a prototypical model with an emerging phase of enhanced quantum fluctuations triggered by first-order magnetic transitions.

cond-mat.str-el

Vision Transformer for Robust Occluded Person Reidentification in Complex Surveillance Scenes

Person re-identification (ReID) in surveillance is challenged by occlusion, viewpoint distortion, and poor image quality. Most existing methods rely on complex modules or perform well only on clear frontal images. We propose Sh-ViT (Shuffling Vision Transformer), a lightweight and robust model for occluded person ReID. Built on ViT-Base, Sh-ViT introduces three components: First, a Shuffle module in the final Transformer layer to break spatial correlations and enhance robustness to occlusion and blur; Second, scenario-adapted augmentation (geometric transforms, erasing, blur, and color adjustment) to simulate surveillance conditions; Third, DeiT-based knowledge distillation to improve learning with limited labels.To support real-world evaluation, we construct the MyTT dataset, containing over 10,000 pedestrians and 30,000+ images from base station inspections, with frequent equipment occlusion and camera variations. Experiments show that Sh-ViT achieves 83.2% Rank-1 and 80.1% mAP on MyTT, outperforming CNN and ViT baselines, and 94.6% Rank-1 and 87.5% mAP on Market1501, surpassing state-of-the-art methods.In summary, Sh-ViT improves robustness to occlusion and blur without external modules, offering a practical solution for surveillance-based personnel monitoring.

cs.CV

Metrics and evaluations for computational and sustainable AI efficiency

The rapid advancement of Artificial Intelligence (AI) has created unprecedented demands for computational power, yet methods for evaluating the performance, efficiency, and environmental impact of deployed models remain fragmented. Current approaches often fail to provide a holistic view, making it difficult to compare and optimise systems across heterogeneous hardware, software stacks, and numeric precisions. To address this gap, we propose a unified and reproducible methodology for AI model inference that integrates computational and environmental metrics under realistic serving conditions. Our framework provides a pragmatic, carbon-aware evaluation by systematically measuring latency and throughput distributions, energy consumption, and location-adjusted carbon emissions, all while maintaining matched accuracy constraints for valid comparisons. We apply this methodology to multi-precision models across diverse hardware platforms, from data-centre accelerators like the GH200 to consumer-level GPUs such as the RTX 4090, running on mainstream software stacks including PyTorch, TensorRT, and ONNX Runtime. By systematically categorising these factors, our work establishes a rigorous benchmarking framework that produces decision-ready Pareto frontiers, clarifying the trade-offs between accuracy, latency, energy, and carbon. The accompanying open-source code enables independent verification and facilitates adoption, empowering researchers and practitioners to make evidence-based decisions for sustainable AI deployment.

cs.PF