arXiv ScienceSearch

arXiv subjects

Xia Xiao

Publications and source records attributed to Xia Xiao.

12 recordsLinked to original sources

Nonuniqueness of Model Potentials and Weak Geodesic Rays

We establish two nonuniqueness phenomena on $\mathbb{CP}^1$. First, distinct normalized positive-mass model potentials can have the same non-pluripolar Monge--Amp\`ere measure. By the contact-set formula, their zero-contact sets agree $\omega_{\mathrm{FS}}$-almost everywhere. Moreover, both have zero Lelong number at every point and identical multiplier ideal sheaves at every positive scale. Second, we construct distinct maximal test curves whose zero-contact sets agree pointwise at every parameter. Their inverse Legendre transforms are distinct bounded weak geodesic rays from zero with the same pointwise right initial tangent.

math.DG

Scaling Reasoning Tokens via RL and Parallel Thinking: Evidence From Competitive Programming

We study how to scale reasoning token budgets for competitive programming through two complementary approaches: training-time reinforcement learning (RL) and test-time parallel thinking. During RL training, we observe an approximately log-linear relationship between validation accuracy and the average number of generated reasoning tokens over successive checkpoints, and show two ways to shift this training trajectory: verification RL warmup raises the starting point, while randomized clipping produces a steeper trend in the observed regime. As scaling single-generation reasoning during RL quickly becomes expensive under full attention, we introduce a multi-round parallel thinking pipeline that distributes the token budget across threads and rounds of generation, verification, and refinement. We train the model end-to-end on this pipeline to match the training objective to the test-time structure. Starting from Seed-OSS-36B, the full system with 16 threads and 16 rounds per thread matches the underlying RL model's oracle pass@16 at pass@1 using 7.6 million tokens per problem on average, and surpasses GPT-5-high on 456 hard competitive programming problems from AetherCode.

cs.CL

On Weighted Twisted K-Energy and Its Applications

We establish the convexity of the weighted twisted Mabuchi K-energy functional along geodesics in the finite energy space $\mathcal{E}^{1,T}(X,\omega)$, covering the case of divisors with mixed cusp and conic singularities. We then prove that coercivity (relative to the complex torus) of this functional is an open condition under cone angle perturbations. This is obtained from a general result of independent interest, which shows the stability of the coercivity under perturbations by certain twist currents. In particular, this yields the openness for the existence for cscK cone metrics and proves that coercivity at the cusp limit implies existence of cscK cone metrics for small cone angles.

math.DG

Scaling up Multi-Turn Off-Policy RL and Multi-Agent Tree Search for LLM Step-Provers

The integration of Large Language Models (LLMs) into automated theorem proving has shown immense promise, yet is fundamentally constrained by challenges in scaling up both training-time reinforcement learning (RL) and inference-time compute. This paper introduces \texttt{BFS-Prover-V2}, a system designed to address this dual scaling problem. We present two primary innovations. The first is a novel multi-turn off-policy RL framework for continually improving the performance of LLM step-prover at training time. This framework, inspired by the principles of AlphaZero, utilizes a multi-stage expert iteration pipeline featuring adaptive tactic-level data filtering and periodic retraining to surmount the performance plateaus that typically curtail long-term RL in LLM-based agents. The second innovation is a planner-enhanced multi-agent search architecture that scales reasoning capabilities at inference time. This architecture employs a general reasoning model as a high-level planner to iteratively decompose complex theorems into a sequence of simpler subgoals. This hierarchical approach substantially reduces the search space, enabling a team of parallel prover agents to collaborate efficiently by leveraging a shared proof cache. We demonstrate that this dual approach to scaling yields state-of-the-art results on established formal mathematics benchmarks. \texttt{BFS-Prover-V2} achieves 95.08\% and 41.4\% on the MiniF2F and ProofNet test sets respectively. While demonstrated in the domain of formal mathematics, the RL and inference techniques presented in this work are of broader interest and may be applied to other domains requiring long-horizon multi-turn reasoning and complex search.

cs.AI

Seed-Coder: Let the Code Model Curate Data for Itself

Code data in large language model (LLM) pretraining is recognized crucial not only for code-related tasks but also for enhancing general intelligence of LLMs. Current open-source LLMs often heavily rely on human effort to produce their code pretraining data, such as employing hand-crafted filtering rules tailored to individual programming languages, or using human-annotated data to train quality filters. However, these approaches are inherently limited in scalability, prone to subjective biases, and costly to extend and maintain across diverse programming languages. To address these challenges, we introduce Seed-Coder, a series of open-source LLMs comprising base, instruct and reasoning models of 8B size, minimizing human involvement in data construction. Our code pretraining data is produced by a model-centric data pipeline, which predominantly leverages LLMs for scoring and filtering code data. The instruct model is further trained via supervised fine-tuning and preference optimization, and the reasoning model leverages Long-Chain-of-Thought (LongCoT) reinforcement learning to improve multi-step code reasoning. Seed-Coder achieves state-of-the-art results among open-source models of similar size and even surpasses some much larger models, demonstrating superior performance in code generation, code completion, code editing, code reasoning, and software engineering tasks.

cs.CL

BFS-Prover: Scalable Best-First Tree Search for LLM-based Automatic Theorem Proving

Recent advancements in large language models (LLMs) have spurred growing interest in automatic theorem proving using Lean4, where effective tree search methods are crucial for navigating the underlying large proof search spaces. While the existing approaches primarily rely on value functions and/or Monte Carlo Tree Search (MCTS), the potential of simpler methods like Best-First Tree Search (BFS) remains underexplored. In this paper, we investigate whether BFS can achieve competitive performance in large-scale theorem proving tasks. We present BFS-Prover, a scalable expert iteration framework, featuring three key innovations. First, we implement strategic data filtering at each expert iteration round, excluding problems solvable via beam search node expansion to focus on harder cases. Second, we improve the sample efficiency of BFS through Direct Preference Optimization (DPO) applied to state-tactic pairs automatically annotated with compiler error feedback, refining the LLM's policy to prioritize productive expansions. Third, we employ length normalization in BFS to encourage exploration of deeper proof paths. BFS-Prover achieves a state-of-the-art score of $72.95\%$ on the MiniF2F test set and therefore challenges the perceived necessity of complex tree search methods, demonstrating that BFS can achieve competitive performance when properly scaled. To facilitate further research and development in this area, we have open-sourced our model at https://huggingface.co/ByteDance-Seed/BFS-Prover-V1-7B.

cs.AI

FullStack Bench: Evaluating LLMs as Full Stack Coders

As the capabilities of code large language models (LLMs) continue to expand, their applications across diverse code intelligence domains are rapidly increasing. However, most existing datasets only evaluate limited application domains. To address this gap, we have developed a comprehensive code evaluation dataset FullStack Bench focusing on full-stack programming, which encompasses a wide range of application domains (e.g., basic programming, data analysis, software engineering, mathematics, and machine learning). Besides, to assess multilingual programming capabilities, in FullStack Bench, we design real-world instructions and corresponding unit test cases from 16 widely-used programming languages to reflect real-world usage scenarios rather than simple translations. Moreover, we also release an effective code sandbox execution tool (i.e., SandboxFusion) supporting various programming languages and packages to evaluate the performance of our FullStack Bench efficiently. Comprehensive experimental results on our FullStack Bench demonstrate the necessity and effectiveness of our FullStack Bench and SandboxFusion.

cs.AI

A high-Q metasurface signal isolator for 1.5T surface coil magnetic resonance imaging on the go

The combination of surface coils and metamaterials remarkably enhance magnetic resonance imaging (MRI) performance for significant local staging flexibility. However, due to the coupling in between, impeded signal-to-noise ratio (SNR) and low-contrast resolution, further hamper the future growth in clinical MRI. In this paper, we propose a high-Q metasurface decoupling isolator fueled by topological LC loops for 1.5T surface coil MRI system, increasing the magnetic field up to fivefold at 63.8 MHz. We have employed a polarization conversion mechanism to effectively eliminate the coupling between the MRI metamaterial and the radio frequency (RF) surface transmitter-receiver coils. Furthermore, a high-Q metasurface isolator was achieved by taking advantage of bound states in the continuum (BIC) for extremely high-field MRI and spectroscopy. An equivalent physical model of the miniaturized metasurface design was put forward through LC circuit analysis. This study opens up a promising route for the easy-to-use and portable surface coil MRI scanners.

physics.med-ph

Field-wise Embedding Size Search via Structural Hard Auxiliary Mask Pruning for Click-Through Rate Prediction

Feature embeddings are one of the most essential steps when training deep learning based Click-Through Rate prediction models, which map high-dimensional sparse features to dense embedding vectors. Classic human-crafted embedding size selection methods are shown to be "sub-optimal" in terms of the trade-off between memory usage and model capacity. The trending methods in Neural Architecture Search (NAS) have demonstrated their efficiency to search for embedding sizes. However, most existing NAS-based works suffer from expensive computational costs, the curse of dimensionality of the search space, and the discrepancy between continuous search space and discrete candidate space. Other works that prune embeddings in an unstructured manner fail to reduce the computational costs explicitly. In this paper, to address those limitations, we propose a novel strategy that searches for the optimal mixed-dimension embedding scheme by structurally pruning a super-net via Hard Auxiliary Mask. Our method aims to directly search candidate models in the discrete space using a simple and efficient gradient-based method. Furthermore, we introduce orthogonal regularity on embedding tables to reduce correlations within embedding columns and enhance representation capacity. Extensive experiments demonstrate it can effectively remove redundant embedding dimensions without great performance loss.

cs.IR

Enhanced Exploration in Neural Feature Selection for Deep Click-Through Rate Prediction Models via Ensemble of Gating Layers

Feature selection has been an essential step in developing industry-scale deep Click-Through Rate (CTR) prediction systems. The goal of neural feature selection (NFS) is to choose a relatively small subset of features with the best explanatory power as a means to remove redundant features and reduce computational cost. Inspired by gradient-based neural architecture search (NAS) and network pruning methods, people have tackled the NFS problem with Gating approach that inserts a set of differentiable binary gates to drop less informative features. The binary gates are optimized along with the network parameters in an efficient end-to-end manner. In this paper, we analyze the gradient-based solution from an exploration-exploitation perspective and use empirical results to show that Gating approach might suffer from insufficient exploration. To improve the exploration capacity of gradient-based solutions, we propose a simple but effective ensemble learning approach, named Ensemble Gating. We choose two public datasets, namely Avazu and Criteo, to evaluate this approach. Our experiments show that, without adding any computational overhead or introducing any hyper-parameter (except the size of the ensemble), our method is able to consistently improve Gating approach and find a better subset of features on the two datasets with three different underlying deep CTR prediction models.

cs.LG

Predicting band gaps and band-edge positions of oxide perovskites using DFT and machine learning

Density functional theory within the local or semilocal density approximations (DFT-LDA/GGA) has become a workhorse in electronic structure theory of solids, being extremely fast and reliable for energetics and structural properties, yet remaining highly inaccurate for predicting band gaps of semiconductors and insulators. Accurate prediction of band gaps using firstprinciples methods is time consuming, requiring hybrid functionals, quasi-particle GW, or quantum Monte Carlo methods. Efficiently correcting DFT-LDA/GGA band gaps and unveiling the main chemical and structural factors involved in this correction is desirable for discovering novel materials in high-throughput calculations. In this direction, we use DFT and machine learning techniques to correct band gaps and band-edge positions of a representative subset of ABO3 perovskite oxides. Relying on results of HSE06 hybrid functional calculations as target values of band gaps, we find a systematic band gap correction of ~1.5 eV for this class of materials, where ~1 eV comes from downward shifting the valence band and ~0.5 eV from uplifting the conduction band. The main chemical and structural factors determining the band gap correction are determined through a feature selection procedure.

cond-mat.mtrl-sci

WiEps: Measurement of Dielectric Property with Commodity WiFi Device -- An application to Ethanol/Water Mixture

WiFi signal has become accessible everywhere, providing high-speed data transmission experience. Besides the communication service, channel state information (CSI) of the WiFi signals is widely employed for numerous Internet of Things (IoT) applications. Recently, most of these applications are based on analysis of the microwave reflections caused by physical movement of the objective. In this paper, a novel contactless wireless sensing technique named WiEps is developed to measure the dielectric properties of the material, exploiting the transmission characteristics of the WiFi signals. In WiEps, the material under test is placed between the transmitter antenna and receiver antenna. A theoretical model is proposed to quantitatively describe the relationship between CSI data and dielectric properties of the material. During the experiment, the phase and amplitude of the transmitted WiFi signals are extracted from the measured CSI data. The parameters of the theoretical model are calculated using measured data from the known materials. Then, WiEps is utilized to estimate the dielectric properties of unknown materials. The proposed technique is first applied to the ethanol/water mixtures. Then, additional liquids are measured for further verification. The estimated permittivities and conductivities show good agreement with the actual values, with the average error of 4.0% and 8.9%, respectively, indicating the efficacy of WiEps. By measuring the dielectric property, this technique is promising to be applied to new IoT applications using ubiquitous WiFi signals, such as food engineering, material manufacturing process monitoring, and security check.

eess.SP