arXiv ScienceSearch

arXiv subjects

Ziyao Xu

Publications and source records attributed to Ziyao Xu.

At least 19 recordsLinked to original sources

A non-conforming finite difference discrete fracture model based on an energy principle

We propose a finite difference discrete fracture model for single-phase flow in fractured porous media. The method is derived from a unified energy principle and is implemented on Cartesian grids that need not conform to the fractures. It introduces no additional degrees of freedom, modifies the underlying finite difference scheme only locally, handles both highly conductive fractures and low-permeability barriers, and naturally preserves symmetry and positive definiteness of the discrete system. Barrier-induced pressure jumps and fracture-induced fluxes can be recovered through inexpensive local post-processing, yielding a sharper representation of the interface effects. Numerical experiments based on manufactured solutions and published $2$D and $3$D benchmark problems demonstrate the accuracy, effectiveness, and flexibility of the method for isolated fractures, complex networks, and anisotropic porous media.

math.NA

An unfitted finite element discrete fracture model for low-permeability barriers via local stiffness matrix modification

Finite element methods are among the most widely used discretizations for flow in porous media, and their discrete fracture models (DFMs) for highly conductive fractures are well-established. Low-permeability barriers, by contrast, have long resisted this framework because they induce pressure discontinuities that continuous elements cannot represent directly. In this paper, we propose a simple extension of the linear finite element method for modeling low-permeability barriers through a closed-form modification of the local stiffness matrix of each barrier-cut element. The method retains exactly the same $H^1$-conforming $P^1$-finite element space, preserves the original sparsity pattern and symmetric positive definiteness, requires no mesh fitting, and coincides with the standard finite element method away from barriers. Moreover, an inexpensive, purely local post-processing step recovers the discontinuous pressure field with a sharp jump at the barrier interface, thereby removing the one-element-wide smearing present in the continuous solution. Convergence studies with manufactured solutions and two- and three-dimensional benchmark problems from the literature confirm the effectiveness of the method.

math.NA

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model

We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters. For SFT data and trajectory construction, we expand the diversity of executable environments, task assets, and agentic scaffolds through real-world deployment and large-scale synthesis. Our RL pipeline applies mixed-mode RLHF over Think and Non-Think responses to improve overall model quality and reduce failure cases, length-controlled reasoning RL to balance accuracy and reasoning efficiency, and agentic RL with outcome and process rewards to stabilize long-horizon training. Extensive evaluations show that Nanbeige4.2-3B outperforms larger models, including Qwen3.5-9B and Gemma4-12B, across diverse agentic benchmarks while remaining competitive on reasoning and alignment tasks. Performance with OpenClaw further supports its use as a compact local personal assistant.

cs.AI

TVD and TVB preservation without TVD time discretization for discontinuous Galerkin methods

Total variation diminishing (TVD) and total variation bounded (TVB) properties are crucial for controlling spurious oscillations in numerical solutions of conservation laws. In the classical Runge--Kutta (RK) discontinuous Galerkin (DG) framework, enforcing these properties is intrinsically tied to TVD time integrators, more commonly known today as strong-stability-preserving (SSP) methods. This reliance imposes severe structural restrictions, including order barriers and incompatibility with various fully discrete DG formulations, ranging from the recent RKDG method with compact stencils (cRKDG) to the widely established Arbitrary DERivative (ADER) DG method. To bypass these constraints, we propose a novel trace-limited corrector framework that preserves the TVD/TVB-in-the-means properties using generic, non-SSP time stepping. Based on Harten's lemma, our key insight is that total variation stability is dictated solely by the cell-average update in the final corrector stage. Consequently, we modify the traces in the numerical fluxes exclusively in the final stage, leaving the intermediate predictor stages unconstrained. This strategy decouples oscillation control from the SSP restriction, accommodates standard RKDG, cRKDG, and ADER-DG predictors, and retains the compactness of the cRKDG framework. We also prove that the limiter does not activate in smooth regions, thereby preserving the underlying accuracy. Finally, numerical experiments are presented to demonstrate the capabilities and robustness of the method.

math.NA

Investigating More Explainable and Partition-Free Compositionality Estimation for LLMs: A Rule-Generation Perspective

Compositional generalization tests are often used to estimate the compositionality of LLMs. However, such tests have the following limitations: (1) they only focus on the output results without considering LLMs' understanding of sample compositionality, resulting in explainability defects; (2) they rely on dataset partition to form the test set with combinations unseen in the training set, suffering from combination leakage issues. In this work, we propose a novel rule-generation perspective for compositionality estimation for LLMs. It requires LLMs to generate a program as rules for dataset mapping and provides estimates of the compositionality of LLMs using complexity-based theory. The perspective addresses the limitations of compositional generalization tests and provides a new way to analyze the compositionality characterization of LLMs. We conduct experiments and analysis of existing advanced LLMs based on this perspective on a string-to-grid task, and find various compositionality characterizations and compositionality deficiencies exhibited by LLMs.

cs.AI

CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution

Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largely unexplored. In open-domain applications visual and textual cues often conflict, requiring models to perform structured reasoning beyond surface-level alignment. We introduce CrossCheck-Bench, a diagnostic benchmark for evaluating contradiction detection in multimodal inputs. The benchmark adopts a hierarchical task framework covering three levels of reasoning complexity and defines seven atomic capabilities essential for resolving cross-modal inconsistencies. CrossCheck-Bench includes 15k question-answer pairs sourced from real-world artifacts with synthetically injected contradictions. The dataset is constructed through a multi-stage annotation pipeline involving more than 450 expert hours to ensure semantic validity and calibrated difficulty across perception, integration, and reasoning. We evaluate 13 state-of-the-art vision-language models and observe a consistent performance drop as tasks shift from perceptual matching to logical contradiction detection. Most models perform well on isolated entity recognition but fail when multiple clues must be synthesized for conflict reasoning. Capability-level analysis further reveals uneven skill acquisition, especially in tasks requiring multi-step inference or rule-based validation. Additional probing shows that conventional prompting strategies such as Chain-of-Thought and Set-of-Mark yield only marginal gains. By contrast, methods that interleave symbolic reasoning with grounded visual processing achieve more stable improvements. These results highlight a persistent bottleneck in multimodal reasoning and suggest new directions for building models capable of robust cross-modal verification.

cs.CL

A Probabilistic Inference Scaling Theory for LLM Self-Correction

Large Language Models (LLMs) have demonstrated the capability to refine their generated answers through self-correction, enabling continuous performance improvement over multiple rounds. However, the mechanisms underlying how and why accuracy evolves during this iterative process remain unexplored. To fill this gap, we propose a probabilistic theory to model the dynamics of accuracy change and explain the performance improvements observed in multi-round self-correction. Through mathematical derivation, we establish that the accuracy after the $t^{th}$ round of self-correction is given by: $Acc_t = Upp - \alpha^t(Upp - Acc_0),$ where $Acc_0$ denotes the initial accuracy, $Upp$ represents the upper bound of accuracy convergence, and $\alpha$ determines the rate of convergence. Based on our theory, these parameters can be calculated and the predicted accuracy curve then can be obtained through only a single round of self-correction. Extensive experiments across diverse models and datasets demonstrate that our theoretical predictions align closely with empirical accuracy curves, validating the effectiveness of the theory. Our work provides a theoretical foundation for understanding LLM self-correction, thus paving the way for further explorations.

cs.CL

A Conservative and Positivity-Preserving Discontinuous Galerkin Method for the Population Balance Equation

We develop a conservative, positivity-preserving discontinuous Galerkin (DG) method for the population balance equation (PBE), which models the distribution of particle numbers across particle sizes due to growth, nucleation, aggregation, and breakage. To ensure number conservation in growth and mass conservation in aggregation and breakage, we design a DG scheme that applies standard treatment for growth and nucleation, and introduces a novel discretization for aggregation and breakage. The birth and death terms are discretized in a symmetric double-integral form, evaluated using a common refinement of the integration domain and carefully selected quadrature rules. Beyond conservation, we focus on preserving the positivity of the number density in aggregation-breakage. Since local mass corresponds to the first moment, the classical Zhang-Shu limiter, which preserves the zeroth moment (cell average), is not directly applicable. We address this by proving the positivity of the first moment on each cell and constructing a moment-conserving limiter that enforces nonnegativity across the domain. To our knowledge, this is the first work to develop a positivity-preserving algorithm that conserves a prescribed moment. Numerical results verify the accuracy, conservation, and robustness of the proposed method.

math.NA

Exponential Time Differencing Runge-Kutta Discontinuous Galerkin (ETD-RKDG) Methods for Nonlinear Degenerate Parabolic Equations

In this paper, we study high-order exponential time differencing Runge-Kutta (ETD-RK) discontinuous Galerkin (DG) methods for nonlinear degenerate parabolic equations. This class of equations exhibits hyperbolic behavior in degenerate regions and parabolic behavior in non-degenerate regions, resulting in sharp wave fronts in the solution profiles and a parabolic-type time-step restriction, $\tau \sim O(h^2)$, for explicit time integration. To address these challenges and solve such equations in complex domains, we employ DG methods with appropriate stabilizing limiters on unstructured meshes to capture the wave fronts and use ETD-RK methods for time integration to resolve the stiffness of parabolic terms. We extract the system's stiffness using the Jacobian matrix of the DG discretization for diffusion terms and adopt a nodal formulation to facilitate its computation. The algorithm is described in detail for two-dimensional triangular meshes. We also conduct a linear stability analysis in one spatial dimension and present computational results on three-dimensional simplex meshes, demonstrating significant improvements in stability and large time-step sizes.

math.NA

Stability and Time-Step Constraints of Exponential Time Differencing Runge--Kutta Discontinuous Galerkin Methods for Advection-Diffusion Equations

In this paper, we investigate the stability and time-step constraints for solving advection-diffusion equations using exponential time differencing (ETD) Runge-Kutta (RK) methods in time and discontinuous Galerkin (DG) methods in space. We demonstrate that the resulting fully discrete scheme is stable when the time-step size is upper bounded by a constant. More specifically, when central fluxes are used for the advection term, the schemes are stable under the time-step constraint tau <= tau_0 * d / a^2, while when upwind fluxes are used, the schemes are stable if tau <= max{tau_0 * d / a^2, c_0 * h / a}. Here, tau is the time-step size, h is the spatial mesh size, and a and d are constants for the advection and diffusion coefficients, respectively. The constant c_0 is the CFL constant for the explicit RK method for the purely advection equation, and tau_0 is a constant that depends on the order of the ETD-RK method. These stability conditions are consistent with those of the implicit-explicit RKDG method. The time-step constraints are rigorously proved for the lowest-order case and are validated through Fourier analysis for higher-order cases. Notably, the constant tau_0 in the fully discrete ETD-RKDG schemes appears to be determined by the stability condition of their semi-discrete (continuous in space, discrete in time) ETD-RK counterparts and is insensitive to the polynomial degree and the specific choice of the DG method. Numerical examples, including problems with nonlinear convection in one and two dimensions, are provided to validate our findings.

math.NA

CiteCheck: Towards Accurate Citation Faithfulness Detection

Citation faithfulness detection is critical for enhancing retrieval-augmented generation (RAG) systems, yet large-scale Chinese datasets for this task are scarce. Existing methods face prohibitive costs due to the need for manually annotated negative samples. To address this, we introduce the first large-scale Chinese dataset CiteCheck for citation faithfulness detection, constructed via a cost-effective approach using two-stage manual annotation. This method balances positive and negative samples while significantly reducing annotation expenses. CiteCheck comprises training and test splits. Experiments demonstrate that: (1) the test samples are highly challenging, with even state-of-the-art LLMs failing to achieve high accuracy; and (2) training data augmented with LLM-generated negative samples enables smaller models to attain strong performance using parameter-efficient fine-tuning. CiteCheck provides a robust foundation for advancing citation faithfulness detection in Chinese RAG systems. The dataset is publicly available to facilitate research.

cs.CL

Investigating the (De)Composition Capabilities of Large Language Models in Natural-to-Formal Language Conversion

To achieve generalized and robust natural-to-formal language conversion (N2F), large language models (LLMs) need to have strong capabilities of decomposition and composition in N2F when faced with an unfamiliar formal language and be able to cope with compositional gaps and counter-intuitive symbolic names. To investigate whether LLMs have this set of basic capabilities in N2F, we propose the DEDC framework. This framework semi-automatically performs sample and task construction, allowing decoupled evaluation of the set of decomposition and composition capabilities of LLMs in N2F. Based on this framework, we evaluate and analyze the most advanced LLMs, and the main findings include that: (1) the LLMs are deficient in both decomposition and composition; (2) the LLMs show a wide coverage of error types that can be attributed to deficiencies in natural language understanding and the learning and use of symbolic systems; (3) compositional gaps and counter-intuitive symbolic names both affect the decomposition and composition of the LLMs. Our work provides a new perspective for investigating the basic capabilities of decomposition and composition of LLMs in N2F. The detailed analysis of deficiencies and attributions can help subsequent improvements of LLMs.

cs.CL

Kimi k1.5: Scaling Reinforcement Learning with LLMs

Language model pretraining with next token prediction has proved effective for scaling compute but is limited to the amount of available training data. Scaling reinforcement learning (RL) unlocks a new axis for the continued improvement of artificial intelligence, with the promise that large language models (LLMs) can scale their training data by learning to explore with rewards. However, prior published work has not produced competitive results. In light of this, we report on the training practice of Kimi k1.5, our latest multi-modal LLM trained with RL, including its RL training techniques, multi-modal data recipes, and infrastructure optimization. Long context scaling and improved policy optimization methods are key ingredients of our approach, which establishes a simplistic, effective RL framework without relying on more complex techniques such as Monte Carlo tree search, value functions, and process reward models. Notably, our system achieves state-of-the-art reasoning performance across multiple benchmarks and modalities -- e.g., 77.5 on AIME, 96.2 on MATH 500, 94-th percentile on Codeforces, 74.9 on MathVista -- matching OpenAI's o1. Moreover, we present effective long2short methods that use long-CoT techniques to improve short-CoT models, yielding state-of-the-art short-CoT reasoning results -- e.g., 60.8 on AIME, 94.6 on MATH500, 47.3 on LiveCodeBench -- outperforming existing short-CoT models such as GPT-4o and Claude Sonnet 3.5 by a large margin (up to +550%).

cs.AI

Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs

Large Language Models (LLMs) can correct their self-generated responses, but a decline in accuracy after self-correction is also witnessed. To have a deeper understanding of self-correction, we endeavor to decompose, evaluate, and analyze the self-correction behaviors of LLMs. By enumerating and analyzing answer correctness before and after self-correction, we decompose the self-correction capability into confidence (being confident to correct answers) and critique (turning wrong answers to correct) capabilities, and propose two metrics from a probabilistic perspective to measure these 2 capabilities, along with another metric for overall self-correction capability evaluation. Based on our decomposition and evaluation metrics, we conduct extensive experiments and draw some empirical conclusions. For example, we find different models can exhibit distinct behaviors: some models are confident while others are more critical. We also find the trade-off between the two capabilities (i.e. improving one can lead to a decline in the other) when manipulating model self-correction behavior by prompts or in-context learning. Further, we find a simple yet efficient strategy to improve self-correction capability by transforming Supervision Fine-Tuning (SFT) data format, and our strategy outperforms vanilla SFT in both capabilities and achieves much higher accuracy after self-correction. Our code will be publicly available on GitHub.

cs.CL

KiGRAS: Kinematic-Driven Generative Model for Realistic Agent Simulation

Trajectory generation is a pivotal task in autonomous driving. Recent studies have introduced the autoregressive paradigm, leveraging the state transition model to approximate future trajectory distributions. This paradigm closely mirrors the real-world trajectory generation process and has achieved notable success. However, its potential is limited by the ineffective representation of realistic trajectories within the redundant state space. To address this limitation, we propose the Kinematic-Driven Generative Model for Realistic Agent Simulation (KiGRAS). Instead of modeling in the state space, KiGRAS factorizes the driving scene into action probability distributions at each time step, providing a compact space to represent realistic driving patterns. By establishing physical causality from actions (cause) to trajectories (effect) through the kinematic model, KiGRAS eliminates massive redundant trajectories. All states derived from actions in the cause space are constrained to be physically feasible. Furthermore, redundant trajectories representing identical action sequences are mapped to the same representation, reflecting their underlying actions. This approach significantly reduces task complexity and ensures physical feasibility. KiGRAS achieves state-of-the-art performance in Waymo's SimAgents Challenge, ranking first on the WOMD leaderboard with significantly fewer parameters than other models. The video documentation is available at \url{https://kigras-mach.github.io/KiGRAS/}.

cs.RO

StreamMOTP: Streaming and Unified Framework for Joint 3D Multi-Object Tracking and Trajectory Prediction

3D multi-object tracking and trajectory prediction are two crucial modules in autonomous driving systems. Generally, the two tasks are handled separately in traditional paradigms and a few methods have started to explore modeling these two tasks in a joint manner recently. However, these approaches suffer from the limitations of single-frame training and inconsistent coordinate representations between tracking and prediction tasks. In this paper, we propose a streaming and unified framework for joint 3D Multi-Object Tracking and trajectory Prediction (StreamMOTP) to address the above challenges. Firstly, we construct the model in a streaming manner and exploit a memory bank to preserve and leverage the long-term latent features for tracked objects more effectively. Secondly, a relative spatio-temporal positional encoding strategy is introduced to bridge the gap of coordinate representations between the two tasks and maintain the pose-invariance for trajectory prediction. Thirdly, we further improve the quality and consistency of predicted trajectories with a dual-stream predictor. We conduct extensive experiments on popular nuSences dataset and the experimental results demonstrate the effectiveness and superiority of StreamMOTP, which outperforms previous methods significantly on both tasks. Furthermore, we also prove that the proposed framework has great potential and advantages in actual applications of autonomous driving.

cs.CV

High-order exponential time differencing multi-resolution alternative finite difference WENO methods for nonlinear degenerate parabolic equations

In this paper, we focus on the finite difference approximation of nonlinear degenerate parabolic equations, a special class of parabolic equations where the viscous term vanishes in certain regions. This vanishing gives rise to additional challenges in capturing sharp fronts, beyond the restrictive CFL conditions commonly encountered with explicit time discretization in parabolic equations. To resolve the sharp front, we adopt the high-order multi-resolution alternative finite difference WENO (A-WENO) methods for the spatial discretization. To alleviate the time step restriction from the nonlinear stiff diffusion terms, we employ the exponential time differencing Runge-Kutta (ETD-RK) methods, a class of efficient and accurate exponential integrators, for the time discretization. However, for highly nonlinear spatial discretizations such as high-order WENO schemes, it is a challenging problem how to efficiently form the linear stiff part in applying the exponential integrators, since direct computation of a Jacobian matrix for high-order WENO discretizations of the nonlinear diffusion terms is very complicated and expensive. Here we propose a novel and effective approach of replacing the exact Jacobian of high-order multi-resolution A-WENO scheme with that of the corresponding high-order linear scheme in the ETD-RK time marching, based on the fact that in smooth regions the nonlinear weights closely approximate the optimal linear weights, while in non-smooth regions the stiff diffusion degenerates. The algorithm is described in detail, and numerous numerical experiments are conducted to demonstrate the effectiveness of such a treatment and the good performance of our method. The stiffness of the nonlinear parabolic partial differential equations (PDEs) is resolved well, and large time-step size computations are achieved.

math.NA

SPOR: A Comprehensive and Practical Evaluation Method for Compositional Generalization in Data-to-Text Generation

Compositional generalization is an important ability of language models and has many different manifestations. For data-to-text generation, previous research on this ability is limited to a single manifestation called Systematicity and lacks consideration of large language models (LLMs), which cannot fully cover practical application scenarios. In this work, we propose SPOR, a comprehensive and practical evaluation method for compositional generalization in data-to-text generation. SPOR includes four aspects of manifestations (Systematicity, Productivity, Order invariance, and Rule learnability) and allows high-quality evaluation without additional manual annotations based on existing datasets. We demonstrate SPOR on two different datasets and evaluate some existing language models including LLMs. We find that the models are deficient in various aspects of the evaluation and need further improvement. Our work shows the necessity for comprehensive research on different manifestations of compositional generalization in data-to-text generation and provides a framework for evaluation.

cs.CL