arXiv ScienceSearch

arXiv subjects

Bing Xu

Publications and source records attributed to Bing Xu.

At least 19 recordsLinked to original sources

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization

Large language model (LLM)-based judges are widely adopted for automated evaluation and reward modeling, yet their judgments are often affected by judgment biases. Accurately evaluating these biases is essential for ensuring the reliability of LLM-based judges. However, existing studies typically investigate limited biases under a single judge formulation, either generative or discriminative, lacking a comprehensive evaluation. To bridge this gap, we propose JudgeBiasBench, a benchmark for systematically quantifying biases in LLM-based judges. JudgeBiasBench defines a taxonomy of judgment biases across 4 dimensions, and constructs bias-augmented evaluation instances through a controlled bias injection pipeline, covering 12 representative bias types. We conduct extensive experiments across both generative and discriminative judges, revealing that current judges exhibit significant and diverse bias patterns that often compromise the reliability of automated evaluation. To mitigate judgment bias, we propose bias-aware training that explicitly incorporates bias-related attributes into the training process, encouraging judges to disentangle task-relevant quality from bias-correlated cues. By adopting reinforcement learning for generative judges and contrastive learning for discriminative judges, our methods effectively reduce judgment biases while largely preserving general evaluation capability.

cs.CL

RAPAC-DP: Response-Aligned Pending-Action Compensation for Diffusion Policies under Delayed Execution

Cloud-side inference gives imitation-learning policies access to greater computational resources, but communication and computation delays can degrade control performance. To compensate for these delays, we propose RAPAC-DP, a response-aligned pending-action compensation framework designed for both diffusion- and flow-based action generators. RAPAC-DP encodes the actions already scheduled for execution before the cloud response arrives into a pending-action sequence that serves as the conditioning input to a parameter-efficient compensation pathway. When delay effects are negligible, bypassing this pathway exactly recovers the frozen base policy. For training, RAPAC-DP constructs delay-conditioned samples from delay-free demonstrations, requiring neither explicit system dynamics nor additional delayed demonstrations. At the largest fixed delay tested on Kinetix, RAPAC-DP retained 81.4% of its overall delay-free performance. At the largest fixed delay tested on each RoboMimic task, it achieved a mean success rate of 0.633 across the three tasks. These results demonstrate the effectiveness of pending-action compensation for cloud-deployed imitation-learning policies.

cs.RO

Observational constraints on fractional holographic dark energy in the light of DESI DR2

Based on the fractional entropy from fractional quantum mechanics, fractional holographic dark energy (FHDE) has been proposed with the Hubble horizon as the IR cutoff (FHDEH). We extend this framework by adopting the future event horizon and the particle horizon as the IR cutoff, proposing the FHDEF and FHDEP models. Using the SN+OHD+DESI DR2 dataset to constrain these models, we find that all three models provide a marginally lower $χ^{2}_{min}$ compared to $Λ$CDM but without significant preference according to AIC and BIC. When CMB distance priors are included, the FHDEH and FHDEP models are strongly ruled out. We further analyze the cosmological evolution for these models, and find that only the FHDEF model predicts nearly identical evolutions of $Ω_{m}$ and $Ω_{de}$ to those of the $Λ$CDM model across cosmic history, but its deceleration parameter $q$ deviate from the $Λ$CDM model in the future, indicating richer late time dynamics beyond the standard $Λ$CDM cosmology.

gr-qc

Impact of Interacting Dark Energy on the Growth of Matter Density Perturbations: Observational Constraints from DESI and Multi-Probe Data

We investigate the impact of a non-gravitational dark sector interaction on the growth of matter density perturbations within both the interacting $w$CDM and the dynamical Chevallier-Polarski-Linder (CPL) scenarios. For $w$CDM model, we develop a parameterization for the growth rate based on a second-order approximation for the growth index $γ$ that explicitly includes the coupling constant $α$. Our analysis reveals a theoretical degeneracy: the coupling induces a correction $Δγ\simeq 1.1α$ in both models, allowing an interacting dark energy model to mimic the growth index predicted by certain modified gravity theories. Then, we confront the models with the latest multi-probe observations, including the Pantheon+ sample of Type Ia supernovae, Baryon Acoustic Oscillation (BAO) data from the Sloan Digital Sky Survey (SDSS) and the second data release (DR2) of the Dark Energy Spectroscopic Instrument (DESI), Cosmic Microwave Background (CMB) measurements, Hubble parameter $H(z)$ data, and redshift-space distortion (RSD) measurements. Our analysis finds that the coupling constant is consistent with zero at approximately the $3σ$ and $2σ$ confidence levels for $w$CDM and CPL models, respectively, showing no definitive statistical evidence for a departure from the standard $Λ$CDM cosmology. The observational constraints strongly disfavor the region of parameter space where interacting dark energy can mimic modified gravity, restricting the growth index to a common approximate interval of $0.53 \lesssim γ\lesssim 0.60$ for both models. This reinforces the growth index as a robust diagnostic for distinguishing between a non-minimal interaction in the dark sector and a genuine modification of gravity with current data.

astro-ph.CO

Comparative Periodogram Analysis of 22 Years of Super-Kamiokande Solar $^{8}\mathrm{B}$ Neutrino Data: Classical, Phase-Based, and Information Theoretic Methods

Solar $^8\mathrm{B}$ neutrinos offer a unique probe of solar interior dynamics and neutrino electromagnetic properties. We present a systematic, multi-method periodogram analysis of the 22-year Super-Kamiokande solar neutrino dataset (1996--2018), comparing nine algorithms. Through hierarchical temporal segmentation, we disentangle astrophysical signals from detector systematics. The Generalized Lomb-Scargle (GLS) method provides the most statistically robust detections by correctly handling heteroscedastic uncertainties, whereas classical Lomb-Scargle systematically underestimates significance. The Lafler--Kinman method generally fails, whereas independent algorithms like MHAOV and PDM1 recover consistent periodicities, providing vital cross-validation. In pre-2001 and SK-I data, seven algorithms provide \textit{weak evidence} ($\ln B > 0$) for a $\sim 38.8$ d periodicity. However, this signal is entirely absent in the highest-statistics SK-IV modified flux data, where the Bayes factor decisively favors the null model ($\ln B \ll -5$), indicating it is a transient feature of the early low-statistics era. Conversely, a $\sim 24.3$ d signal in post-2001 raw flux is decisively rejected by the Bayesian framework and vanishes in modified flux, confirming its seasonal systematic origin. Furthermore, no evidence is found for an $\sim 11$-year solar cycle modulation, yielding a stringent amplitude upper limit of $<0.2\%$ of the mean flux. By highlighting the stark contrast between frequentist significance and Bayesian model selection ($\ln B$) in low signal-to-noise regimes, we establish a rigorous, multi-metric best-practice framework for periodicity searches. This work provides a direct methodological blueprint for next-generation observatories like Hyper-Kamiokande and JUNO.

astro-ph.HE

Ascent and descent of bounded linear operators

Let $\mathcal B(\mathcal X)$ be the algebra of all bounded linear operators on a real or complex Banach space $\mathcal{X}$ with $\dim\mathcal X \ge 3$. In this paper, we first explore the ascent (descent) of upper triangular block operator matrices and certain special algebraic operators, and then establish characterizations for the ascent (descent) of rank-one and rank-two operators. Based on these results, we characterize features for some special operators by the ascent (descent) of Jordan products. As an application, we give the structure of all maps with range containing all bounded operators of rank at most three preserving the ascent (descent) of operator Jordan product on $\mathcal B(\mathcal X)$.

math.FA

A Generalizable Light Transport 3D Embedding for Global Illumination

Global illumination (GI) is essential for realistic rendering but remains computationally expensive due to the complexity of simulating indirect light transport. Recent neural methods have mainly relied on per-scene optimization, sometimes extended to handle changes in camera or geometry. Efforts toward cross-scene generalization have largely stayed in 2D screen space, such as neural denoising or G-buffer based GI prediction, which often suffer from view inconsistency and limited spatial understanding. We propose a generalizable 3D light transport embedding that approximates global illumination directly from 3D scene configurations, without using rasterized or path-traced cues. Each scene is represented as a point cloud with geometric and material features. A scalable transformer models global point-to-point interactions to encode these features into neural primitives. At render time, each query point retrieves nearby primitives via nearest-neighbor search and aggregates their latent features through cross-attention to predict the desired rendering quantity. We demonstrate results on diffuse global illumination prediction across diverse indoor scenes with varying layouts, geometry, and materials. The embedding trained for irradiance estimation can be quickly adapted to new rendering tasks with limited fine-tuning. We also present preliminary results for spatial-directional radiance field estimation for glossy materials and show how the normalized field can accelerate unbiased path guiding. This approach highlights a path toward integrating learned priors into rendering pipelines without explicit ray-traced illumination cues.

cs.GR

8DNA: 8D Neural Asset Light Transport by Distribution Learning

High-fidelity 3D assets exhibit intriguing global illumination effects like subsurface scattering, glossy interreflections, and fine-scale fiber scatterings, which often involve long scattering paths that are expensive to simulate. We introduce 8D neural assets (8DNA) to pre-bake these light transport effects into neural representations. Unlike prior methods that assume far-field lighting and precompute light transport into 6D functions, 8DNA learns the full 8D light transport, enabling accurate rendering under near-field illumination. Our training leverages a distribution-learning formulation that learns light transport from forward path-traced samples, which produces less optimization variance with lower training budget than the prior regression-based approaches. Experiments show our 8DNA rendering closely matches path-traced results under various scene configurations, yet it achieves improved variance reduction and fast inference speeds on challenging assets.

cs.GR

Enhancing Online Recruitment with Category-Aware MoE and LLM-based Data Augmentation

Person-Job Fit (PJF) is a critical component for online recruitment. Existing approaches face several challenges, particularly in handling low-quality job descriptions and similar candidate-job pairs, which impair model performance. To address these challenges, this paper proposes a large language model (LLM) based method with two novel techniques: (1) LLM-based data augmentation, which polishes and rewrites low-quality job descriptions by leveraging chain-of-thought (COT) prompts, and (2) category-aware Mixture of Experts (MoE) that assists in identifying similar candidate-job pairs. This MoE module incorporates category embeddings to dynamically assign weights to the experts and learns more distinguishable patterns for similar candidate-job pairs. We perform offline evaluations and online A/B tests on our recruitment platform. Our method relatively surpasses existing methods by 2.40% in AUC and 7.46% in GAUC, and boosts click-through conversion rate (CTCVR) by 19.4% in online tests, saving millions of CNY in external headhunting expenses.

cs.AI

AVO: Agentic Variation Operators for Autonomous Evolutionary Search

Agentic Variation Operators (AVO) are a new family of evolutionary variation operators that replace the fixed mutation, crossover, and hand-designed heuristics of classical evolutionary search with autonomous coding agents. Rather than confining a language model to candidate generation within a prescribed pipeline, AVO instantiates variation as a self-directed agent loop that can consult the current lineage, a domain-specific knowledge base, and execution feedback to propose, repair, critique, and verify implementation edits. We evaluate AVO on attention, among the most aggressively optimized kernel targets in AI, on NVIDIA Blackwell (B200) GPUs. Over 7 days of continuous autonomous evolution on multi-head attention, AVO discovers kernels that outperform cuDNN by up to 3.5% and FlashAttention-4 by up to 10.5% across the evaluated configurations. The discovered optimizations transfer readily to grouped-query attention, requiring only 30 minutes of additional autonomous adaptation and yielding gains of up to 7.0% over cuDNN and 9.3% over FlashAttention-4. Together, these results show that agentic variation operators move beyond prior LLM-in-the-loop evolutionary pipelines by elevating the agent from candidate generator to variation operator, and can discover performance-critical micro-architectural optimizations that produce kernels surpassing state-of-the-art expert-engineered attention implementations on today's most advanced GPU hardware.

cs.LG

Long-form RewardBench: Evaluating Reward Models for Long-form Generation

The widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models in various domains and scenarios. However, a significant gap remains in assessing reward models for long-form generation, despite its critical role in real-world applications. To bridge this, we introduce Long-form RewardBench, the first reward modeling testbed specifically designed for long-form generation. Our benchmark encompasses five key subtasks: QA, RAG, Chat, Writing, and Reasoning. We collected instruction and preference data through a meticulously designed multi-stage data collection process, and conducted extensive experiments on 20+ mainstream reward models, including both classifiers and generative models. Our findings reveal that current models still lack long-form reward modeling capabilities. Furthermore, we designed a novel Long-form Needle-in-a-Haystack Test, which revealed a correlation between reward modeling performance and the error's position within a response, as well as the overall response length, with distinct characteristics observed between classification and generative models. Finally, we demonstrate that classifiers exhibit better generalizability compared to generative models trained on the same data. As the first benchmark for long-form reward modeling, this work aims to offer a robust platform for visualizing progress in this crucial area.

cs.CL

Orbital-Selective Spin-Orbit Mott Insulator in Fractional Valence Iridate La$_3$Ir$_3$O$_{11}$

The combination of strong spin-orbit coupling and Coulomb interactions makes the $5d$ iridates a unique platform for realizing novel correlated electronic states. Here, utilizing infrared spectroscopy, we demonstrate that a robust Mott insulating state persists in the $1/3$-hole self-doped system La$_3$Ir$_3$O$_{11}$, evidenced by the collapse of the Drude response and the emergence of sharp excitations across the Mott gap. Our theoretical calculations reveal that the insulating behavior arises from the cooperative interplay of structural distortions, spin-orbit coupling, and Coulomb interactions. Specifically, octahedral distortion and Ir-Ir dimerization split the $t_{2g}$ orbitals, driving the $J_{\mathrm{eff}} = 1/2$ bands toward half-filling while keeping the $J_{\mathrm{eff}} = 3/2$ bands away from it. Consequently, electron correlations induce an orbital-selective Mott transition in the $J_{\mathrm{eff}} = 1/2$ bands, whereas a band-insulating gap develops in the $J_{\mathrm{eff}} = 3/2$ bands, thereby stabilizing the unconventional insulating state in La$_3$Ir$_3$O$_{11}$. These findings provide new insights into the design and understanding of the insulating ground state of spin-orbit-coupled iridates.

cond-mat.str-el

DSFC-Net: A Dual-Encoder Spatial and Frequency Co-Awareness Network for Rural Road Extraction

Accurate extraction of rural roads from high-resolution remote sensing imagery is essential for infrastructure planning and sustainable development. However, this task presents unique challenges in rural settings due to several factors. These include high intra-class variability and low inter-class separability from diverse surface materials, frequent vegetation occlusions that disrupt spatial continuity, and narrow road widths that exacerbate detection difficulties. Existing methods, primarily optimized for structured urban environments, often underperform in these scenarios as they overlook such distinctive characteristics. To address these challenges, we propose DSFC-Net, a dual-encoder framework that synergistically fuses spatial and frequency-domain information. Specifically, a CNN branch is employed to capture fine-grained local road boundaries and short-range continuity, while a novel Spatial-Frequency Hybrid Transformer (SFT) is introduced to robustly model global topological dependencies against vegetation occlusions. Distinct from standard attention mechanisms that suffer from frequency bias, the SFT incorporates a Cross-Frequency Interaction Attention (CFIA) module that explicitly decouples high- and low-frequency information via a Laplacian Pyramid strategy. This design enables the dynamic interaction between spatial details and frequency-aware global contexts, effectively preserving the connectivity of narrow roads. Furthermore, a Channel Feature Fusion Module (CFFM) is proposed to bridge the two branches by adaptively recalibrating channel-wise feature responses, seamlessly integrating local textures with global semantics for accurate segmentation. Comprehensive experiments on the WHU-RuR+, DeepGlobe, and Massachusetts datasets validate the superiority of DSFC-Net over state-of-the-art approaches.

cs.CV

An Extended VIIRS-like Artificial Nighttime Light Data Reconstruction (1986-2024)

Artificial Night-Time Light (NTL) remote sensing is a vital proxy for quantifying the intensity and spatial distribution of human activities. Although the NPP-VIIRS sensor provides high-quality NTL observations, its temporal coverage, which begins in 2012, restricts long-term time-series studies that extend to earlier periods. Current extended VIIRS-like NTL data products suffer from two significant shortcomings: the underestimation of light intensity and the omission of structural details. To overcome these limitations, we present the Extended VIIRS-like Artificial Nighttime Light (EVAL) dataset, a new annual NTL dataset for China spanning from 1986 to 2024. This dataset was generated using a novel two-stage deep learning model designed to address the aforementioned shortcomings. The model first constructs an initial estimate and subsequently refines fine-grained structural details using high-resolution impervious surface data as guidance. Quantitative evaluations demonstrate that EVAL significantly outperforms state-of-the-art products, exhibiting superior temporal consistency and a stronger correlation with socioeconomic indicators.

cs.CV

VibeTensor: System Software for Deep Learning, Fully Generated by AI Agents

VIBETENSOR is an open-source research system software stack for deep learning, generated by LLM-powered coding agents under high-level human guidance. In this paper, "fully generated" refers to code provenance: implementation changes were produced and applied as agent-proposed diffs; validation relied on agent-run builds, tests, and differential checks, without per-change manual diff review. It implements a PyTorch-style eager tensor library with a C++20 core (CPU+CUDA), a torch-like Python overlay via nanobind, and an experimental Node.js/TypeScript interface. Unlike thin bindings, VIBETENSOR includes its own tensor/storage system, schema-lite dispatcher, reverse-mode autograd, CUDA runtime (streams/events/graphs), a stream-ordered caching allocator with diagnostics, and a stable C ABI for dynamically loaded operator plugins. We view this release as a milestone for AI-assisted software engineering: it shows coding agents can generate a coherent deep learning runtime spanning language bindings down to CUDA memory management, validated primarily by builds and tests. We describe the architecture, summarize the workflow used to produce and validate the system, and evaluate the artifact. We report repository scale and test-suite composition, and summarize reproducible microbenchmarks from an accompanying AI-generated kernel suite, including fused attention versus PyTorch SDPA/FlashAttention. We also report end-to-end training sanity checks on 3 small workloads (sequence reversal, ViT, miniGPT) on NVIDIA H100 (Hopper, SM90) and Blackwell-class GPUs; multi-GPU results are Blackwell-only and use an optional CUTLASS-based ring-allreduce plugin gated on CUDA 13+ and sm103a toolchain support. Finally, we discuss failure modes in generated system software, including a "Frankenstein" composition effect where locally correct subsystems interact to yield globally suboptimal performance.

cs.SE

Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory

The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model capabilities. This paper provides a critical analysis of benchmark effectiveness, examining mainstream prominent LLM benchmarks using results from diverse models. We first propose Pseudo-Siamese Network for Item Response Theory (PSN-IRT), an enhanced Item Response Theory framework that incorporates a rich set of item parameters within an IRT-grounded architecture. PSN-IRT can be utilized for accurate and reliable estimations of item characteristics and model abilities. Based on PSN-IRT, we conduct extensive analysis on 11 LLM benchmarks comprising 41,871 items, revealing significant and varied shortcomings in their measurement quality. Furthermore, we demonstrate that leveraging PSN-IRT is able to construct smaller benchmarks while maintaining stronger alignment with human preference.

cs.CL

Highly Anisotropic Charge Dynamics and Spectral Weight Redistribution in the Trilayer Nickelate La$_{4}$Ni$_{3}$O$_{10}$

We study the $ab$-plane and $c$-axis charge dynamics of La$_{4}$Ni$_{3}$O$_{10}$ using optical spectroscopy. While a pronounced Drude profile, i.e. metallic response, is observed in the $ab$-plane optical conductivity $σ_{1}^{ab}(ω)$, the $c$-axis optical spectra $σ_{1}^{c}(ω)$ exhibit semiconducting behavior. The zero-frequency extrapolation of the optical conductivity $σ_{1}(ω\rightarrow 0) \equiv 1/ρ_{\text{dc}}$ gives a resistivity anisotropy of $ρ_{c}/ρ_{ab} \simeq 366$ at 300~K for La$_{4}$Ni$_{3}$O$_{10}$, which is much larger than the values in iron-based superconductors but comparable to those in high-$T_{c}$ cuprates. The interband response is also highly anisotropic, showing salient orbital selectivity for light polarized in the $ab$ plane and along the $c$ axis. The interband-transition peaks in both $σ_{1}^{ab}(ω)$ and $σ_{1}^{c}(ω)$ are located at lower energies compared to density-functional-theory predictions, signifying considerable electronic correlations. By investigating the spectral weight transfer, we find that in the pristine phase, Coulomb correlations have a marked impact on the charge dynamics of \LNO, whereas in the density-wave state, a gap opens with the Ni-$d_{z^{2}}$ orbital being involved.

cond-mat.supr-con

ELPO: Ensemble Learning Based Prompt Optimization for Large Language Models

The remarkable performance of Large Language Models (LLMs) highly relies on crafted prompts. However, manual prompt engineering is a laborious process, creating a core bottleneck for practical application of LLMs. This phenomenon has led to the emergence of a new research area known as Automatic Prompt Optimization (APO), which develops rapidly in recent years. Existing APO methods such as those based on evolutionary algorithms or trial-and-error approaches realize an efficient and accurate prompt optimization to some extent. However, those researches focus on a single model or algorithm for the generation strategy and optimization process, which limits their performance when handling complex tasks. To address this, we propose a novel framework called Ensemble Learning based Prompt Optimization (ELPO) to achieve more accurate and robust results. Motivated by the idea of ensemble learning, ELPO conducts voting mechanism and introduces shared generation strategies along with different search methods for searching superior prompts. Moreover, ELPO creatively presents more efficient algorithms for the prompt generation and search process. Experimental results demonstrate that ELPO outperforms state-of-the-art prompt optimization methods across different tasks, e.g., improving F1 score by 7.6 on ArSarcasm dataset.

cs.CL