arXiv ScienceSearch

arXiv subjects

Yanbo Wang

Publications and source records attributed to Yanbo Wang.

At least 19 recordsLinked to original sources

DTD-VAE: Disentangled Temporal Dependencies VAE for Credit Risk Prediction

Evaluating customer creditworthiness is crucial for retail banking operations, as it impacts marketing strategies, customer relationship management, and credit risk control. Traditional methods often struggle to capture complex temporal dependencies and extract pertinent information from customer data, crucial for accurate risk assessment. Specifically, they fail to differentiate between temporal patterns indicative of credit risk and those reflecting general customer behavior or preferences, leading to suboptimal risk predictions. In this study, we introduce the Disentangled Temporal Dependencies Variational Autoencoder (DTD-VAE), an advancement over conventional VAE, designed to disentangle temporal dependencies and distinguish credit risk-related features from past customer preferences. The feature inference module of the DTD-VAE incorporates an autoregressive temporal dependency learning mechanism that adeptly captures the temporal dependencies among latent variables, enriching the model's comprehension of the inherent data structure. Furthermore, the feature generative module utilizes an element-wise gating mechanism that assigns independent weights to each dimension of the expert models, enabling a finer-grained disentanglement of latent variables, particularly those relevant to credit risk prediction. Extensive experiments on six real-world datasets demonstrate that the proposed framework consistently outperforms existing methods, achieving performance gains of 3.2%-4.86% in ROC-AUC and 6.41%-9.71% in Accuracy Ratio.

q-fin.RM

Unstable Manifolds of Stratified Euler Equations

We consider a spectrally unstable steady state $(\rho_0,v_0)$ of the incompressible stratified Euler equations on a class of $d$-dimensional domains. Assuming that the linearized equation admits an exponential dichotomy with a reasonably large spectral gap relative to the maximal Lyapunov exponent of the background steady flow $v_0$, we construct the local stable and unstable manifolds of $(\rho_0,v_0)$. The proof is based on the Lyapunov--Perron method after reformulating the Euler equation as an ODE on the infinite-dimensional manifold of volume-preserving Lagrangian maps, with the density treated as a frozen Lagrangian parameter as well as the weight in the $L^2$ metric. We also discuss some applications to two-dimensional steady flows.

math.AP

Data-DPO: Direct Preference Optimization for Target Model Data Selection in LLM Post-Training

Data selection in supervised fine-tuning aims to select a small set of effective samples from large-scale candidate data, reducing training cost while preserving model performance. However, existing methods usually treat data value as a relatively static property, and pay limited attention to the compatibility between data and the capability distribution of the target model. To address this issue, we propose Data-DPO, a target model-oriented SFT data selection method. Data-DPO observes the local training feedback of the target model on different samples through one-step probing, transforms activation differences among samples into pairwise data preferences, and trains a lightweight reward model to learn target-model-aware data preferences. In the final selection stage, Data-DPO further combines target model preference, external quality scores, and marginal diversity to construct a more stable and effective training subset. Experimental results on Vision-Flan and LLaVA-CoT show that Data-DPO consistently outperforms existing data selection baselines under multiple data budgets and stably surpasses full data training performance.

cs.LG

Hierarchical Data Selection via Manifold Coverage and Sparse Feature Coverage in LLM Post-training

As supervised fine-tuning data continues to scale, selecting high-value subsets from large candidate pools is crucial for reducing training cost and improving model performance. Existing methods often measure diversity directly in the original embedding space, where geometric metrics entangle dominant semantic directions, fine-grained supervision differences, and local noise. We address this limitation by formulating data selection as a coarse-to-fine hierarchical coverage problem and propose MASS. MASS learns low-dimensional principal manifold coordinates with a dense autoencoder for coarse semantic grouping, and then performs quality-aware sparse feature coverage within each group using a TopK sparse autoencoder. Experiments on Vision Flan and LLaVA-CoT show that MASS consistently outperforms strong data selection baselines across multiple budgets, and in several settings matches or surpasses full data training with only a small subset of data.

cs.LG

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may reach new answers, or produce answers that were already within reach. Aggregate scores do not distinguish these changes question by question. We establish a question-level audit under fixed budgets, temperatures, and answer formats. A question is realized when the default deployment procedure produces the correct answer. A question is reachable when a specified probe finds that answer within a fixed budget. We first test whether inference-time layer routing can expand reachability. Under a matched budget, random routes match or exceed structured search in all 43 model and task settings. Answer-blind procedures retain almost none of this gain, which instead requires access to the correct answer. We then ask why reachable answers sometimes fail to appear. Across six cases spanning 0.5B to 31B, silencing one identified MLP block repairs 68 to 92 percent of a predefined failure set. We next test whether training closes the gap by expanding reachability. In five of six matched evaluations, deployed performance rises while the reachable ceiling remains flat or falls. For DAPO, the deployed score rises by 14.7 points while the reachable ceiling falls by 13.3 points. Across the settings we audit, realization and reachability therefore do not always change together. Claims of capability expansion should report both realized performance and reachability under matched evaluation conditions. Code is available at https://github.com/LiZaiyuan0619/reachability-not-realization

cs.AI

Disagree to Accelerate: Closing the Loop on Diffusion Feature Forecasts

Training-free feature forecasting accelerates diffusion sampling by predicting features at skipped denoising steps. Recent work has mainly focused on designing stronger forecasters. Yet forecast error varies sharply across steps, and open-loop caches trust the forecast in full at every skipped step. This fixed trust is what breaks as acceleration turns aggressive. The missing question is not only how to forecast better, but when and how much to trust a forecast. We show that reliability can be observed from the cache itself. Two forecasts agree where the feature trajectory is smooth, and they diverge where prediction turns hard. Their disagreement is a cheap runtime signal, and it costs no extra denoiser evaluation. Based on this signal, we introduce RACER, a training-free closed-loop controller with two responses. It continuously shrinks uncertain forecasts toward the last computed feature. At the riskiest steps, RACER refreshes the feature and repays the added evaluation by skipping a later scheduled one. We derive a deterministic error bound for the shrinkage and empirically evaluate its validity and tightness across acceleration regimes. At the same number of denoiser evaluations, RACER improves the strongest open-loop baseline across SD3.5-Large, FLUX.1-dev, Wan2.1-14B, and HunyuanVideo on DrawBench, VBench, and COCO. On SD3.5, we further show that RACER samples faster at equal quality. RACER generalizes across forecasting designs as well. For example, it recovers much of the quality lost on a Taylor base. These results show that reliable diffusion acceleration also depends on how forecasts are used. Code is available at https://github.com/LiZaiyuan0619/RACER

cs.LG

Connecting Microseismicity to Lithology via a Model of Slip Avalanches

Fluid injection into the earth's crust can induce small and frequent earthquakes in the subsurface. Predicting their sizes and temporal occurrences via statistical analysis is crucial for safe operations in unconventional oil and gas recovery, enhanced geothermal systems, and geologic carbon storage. Here we show that a simple micromechanical model of slip avalanches in slowly deforming solids predicts the slip statistics observed over drastically different spatial scales, namely meter-scale microseismic observations and nanometer- to micrometer-scale nanoindentation experiments can be described with this model. Microseismic catalogs extracted from high-pressure fluid injection operations into geological basins with various lithologies and nanoindentation experiments on shale across a wide range of temperatures and mineral compositions yield statistics consistent with model predictions. This universality across materials, temperatures, and scales is consistent with the prediction that the slip statistics result from only a few basic properties. Previously debated deviations of the statistics in layered sedimentary formations are explained by finite-size and stress-integrative effects resulting from mechanically weak bedding planes. The slip statistics therefore provide important information about the structure and scales of the bedding planes. Conversely, the basin structure can also be used to predict the probability distribution for the sizes of triggered microseismic events.

cond-mat.other

No Time Like the Present: Agentic Test-Time Training for LLM Agents

LLM agents often degrade over long episodes: as trajectories grow, they revisit explored states, repeat failed actions, and lose strategies that previously worked. Test-time training (TTT) offers a way to adapt model weights to the evolving task state, but existing LLM TTT methods largely adapt once to a fixed input. We study continuous TTT in multi-turn agent episodes, where each update changes the policy that generates later training text. This creates a self-training loop that helps when new trajectory information appears, but can amplify drift when the agent gets stuck and repeatedly trains on similar text. We find that update-text repetition distinguishes these regimes and introduce Agentic Test-Time Training (aTTT), a token-level reweighting method that downweights the loss on tokens appearing in repeated $n$-grams from prior updates while leaving novel tokens fully weighted. To run such updates inside live episodes, we build a concurrent serving system using vLLM's runtime LoRA API, limiting overhead to 1.9$\times$ the no-TTT cost. aTTT improves success by up to 5.0 points on ALFWorld and 4.9 points on SWE-bench Lite. The gains concentrate where models already have task competence but drift over long trajectories, suggesting that aTTT mainly preserves existing competence rather than teaching new abilities.

cs.LG

MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion

Streaming zero-shot voice conversion (VC) has become increasingly popular due to its potential for real-time applications. The recently proposed MeanVC achieves lightweight streaming zero-shot VC, but it has several limitations: its chunk-wise autoregressive denoising doubles the effective training sequence length, conversion quality degrades under small-chunk settings, and its timbre encoder directly relies on reference mel-spectrograms, making it sensitive to reference audio quality. To address these limitations we propose MeanVC 2. We introduce future-receptive chunking (FRC), which explicitly schedules past and future receptive fields across diffusion transformer decoder layers and removes clean-chunk teacher forcing. By incorporating bounded future context, FRC enables stable conversion with a 40 ms chunk size. We further introduce a universal timbre token encoder, which constructs a timbre representation from a global speaker embedding and retrieves fine-grained timbre cues via cross-attention, improving robustness to low-quality references and enhancing zero-shot speaker similarity. Experimental results show that MeanVC 2 significantly outperforms MeanVC, while reducing latency from 211 ms to 110 ms. Audio samples are publicly available. The source code will be publicly released.

eess.AS

Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing

With the wide adoption of Multimodal Models (MMs) in real-world scenarios, it is significant to efficiently train emerging MMs that exhibit increasingly complex module architectures. For MM deployment, existing works allocate a GPU to only one MM module in a temporal-multiplexing manner; this compromises training efficiency because a single module often fails to achieve high GPU utilization. To improve GPU utilization and enable efficient MM training, we propose deploying MMs in a temporal-spatial multiplexing manner, allowing multiple MM modules to colocate on a GPU with well-controlled resource quotas. In this paper, we propose Apollo, an efficient MM training system that applies temporal-spatial multiplexing. We first develop a flexible and lightweight execution engine that supports MM training with arbitrary resource quotas, and then build a comprehensive and accurate performance model to estimate module execution time under different allocation plans. With the performance model, we further adopt effective heuristics to derive high-quality MM deployment plans efficiently. Testbed experiments confirm that Apollo effectively improves the training efficiency of popular MMs, with a training speedup of up to 1.31x.

cs.DC

SAGE: A Self-Evolving Agentic Graph-Memory Engine for Structure-Aware Associative Memory

Long-term memory is becoming a central bottleneck for language agents. Exsting RAG and GraphRAG systems largely treat memory graphs as static retrieval middleware, which limits their ability to recover complete evidence chains from partial cues, exploit reusable graph-structrual roles, and improve the memory itself through downstream feedback. We introduce SAGE, a Self-evolving Agentic Graph-memory Engine that models graph memory as a dynamic long-term memory substrate. SAGE couples two roles: a memory writer that incrementally constucts structured graph memory from interaction histories, and a Graph Foundation Model-based memory reader to perform retrieval and provide feedback to the memory writer. We provide rigorooous theoretical annalyses supporting the framework. Across multi-hop QA, open-domain retireval, domain-specific review QA, and long-term agent-memory benchmarks, SAGE improves evidence recovery, answer grounding, and retrieval efficiency: after two self-evolution rounds, it achieves the best average rank on multi-hop QA; in zero-shot open-domain transfer, it reaches 82.5/91.6 Recall@2/5 on NQ. Further results on LongMemEval and HaluMem show that traning and reader-writer feedback improve multiple long-term memory and hallucination-diagnostic metrics, suggesting that self-evolving, structure-aware graph memory is a promising foundation for robust long-horizon language agents.

cs.AI

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents

While Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for closed-ended tasks, extending it to open-ended social language games via self-play reveals a critical issue: evolution impasse. Due to the vast strategy space, language agents frequently converge to homogenized behaviors, leading to deterministic match outcomes that eliminate the gradient signals necessary for policy evolution. To tackle this issue, we propose Dual-scale Evolutionary Policy Training (DEPT) for social language games. DEPT introduces a time-scaled evolutionary perception mechanism that detects impasse by quantifying dual-scale value baseline divergence alongside match entropy. Upon perceiving the collapse, it then activates asymmetric advantage reshaping to dynamically modulate the optimization landscape for intervention. Thus, our method effectively restores gradient signals and enforces sustained strategic exploration. Extensive experiments on multiple social language games demonstrate that DEPT outperforms strong baselines, avoiding policy degeneration and driving the continuous evolution of social language agents.

cs.CL

Position: How can Graphs Help Large Language Models?

With the rapid advancement of large language models (LLMs), classic graph learning tasks have greatly benefited from LLMs, including improved encoding of textual features, more efficient construction of graphs from text, and enhanced reasoning over knowledge graphs. In this paper, we ask a complementary question: How can graphs help LLMs? We address this question from three perspectives: 1) graphs provide an up-to-date knowledge source that helps reduce LLM hallucinations, 2) graph-based prompting techniques-such as Chain-of-Thought (CoT), Tree-of-Thought (ToT), and Graph-of-Thought (GoT)-enhance LLM reasoning capabilities, and 3) integrating graphs into LLMs improves their understanding of structured data, expanding their applicability to domains such as e-commerce, code, and relational databases (RDBs). We further outlook some future directions including designing sparse LLM architectures based on graphs and brain-inspired memory systems.

cs.AI

APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation

Privacy policies are essential for users to understand how service providers handle their personal data. However, these documents are often long and complex, as well as filled with technobabble and legalese, causing users to unknowingly accept terms that may even contradict the law. While summarizing and interpreting these privacy policies is crucial, there is a lack of high-quality English parallel corpus optimized for legal clarity and readability. To address this issue, we introduce APPSI-139, a high-quality English privacy policy corpus meticulously annotated by domain experts, specifically designed for summarization and interpretation tasks. The corpus includes 139 English privacy policies, 15,692 rewritten parallel corpora, and 36,351 fine-grained annotation labels across 11 data practice categories. Concurrently, we propose TCSI-pp-V2, a hybrid privacy policy summarization and interpretation framework that employs an alternating training strategy and coordinates multiple expert modules to effectively balance computational efficiency and accuracy. Experimental results show that the hybrid summarization system built on APPSI-139 corpus and the TCSI-pp-V2 framework outperform large language models, such as GPT-4o and LLaMA-3-70B, in terms of readability and reliability. The source code and dataset are available at https://github.com/EnlightenedAI/APPSI-139.

cs.CL

Selecting the optimal Parameters Results in Double Interpolation: Double AFD

Let $f$ belong to the Hardy space $H^2(\mathbb{D})$ of the unit disc, and $e_a$ the normalized Szeg\"o (reproducing) kernel of $H^2(\mathbb{D}).$ It is well known that, due to the reproducing kernel property, for any distinct $n$ points $a_1,\cdots,a_n$ in $\mathbb{D}$ the orthogonal projection of $f$ into ${\rm span}\{e_{a_1},\cdots,e_{a_n}\},$ denoted as $P_{{\rm span}\{e_{a_1},\cdots,e_{a_n}\}}(f),$ interpolates $f$ at the points $a_k$'s. The present study further proves that if the $a_k$'s are optimally selected according to certain energy matching pursuit principle, then $P_{{\rm span}\{e_{a_1},\cdots,e_{a_n}\}}(f)$ double interpolates $f$ at the points $a_k$'s, or order $m=2$ interpolation, that is, \[ P_{{\rm span}\{e_{a_1},\cdots,e_{a_n}\}}(f)(a_k)=f(a_k), \quad {\rm and}\quad P_{{\rm span}\{e_{a_1},\cdots,e_{a_n}\}}'(f)(a_k)=f'(a_k),\quad k=1,\cdots,n.\] With the accordingly newly defined double Takenaka-Malmquist system, the norm convergence for $n\to \infty,$ the $n$-best approximation for $n$ being fixed, and the related boundary function interpolation are studied. The such generated new sparse representation, named as double AFD, is shown to outperform the classical AFD. Pointwise interpolations for orders $m>2,$ meaning to simultaneously interpolates all functions $f,f',\cdots,f^{(m-1)}$ at a set of $a_k$'s are, additionally, discussed. For the Hardy space of the upper-half complex plane there exists a counterpart theory.

math.CV

Galactic Diffuse Gamma-Ray and Neutrino Emission from Cosmic-Ray Interactions in Stellar Atmospheres

The Galactic diffuse gamma-ray emission is conventionally modeled as the product of cosmic-ray interactions with the interstellar medium. However, the cumulative contribution of stellar atmospheres acting as hadronic interaction targets remains an unexplored multi-messenger background. In this work, we present the first systematic evaluation of this stellar diffuse emission by coupling MESA stellar evolution profiles and magnetic-field-modulated cosmic-ray transport with a 3D Galactic population synthesis framework. We find that the cumulative stellar contribution to the Galactic diffuse gamma-ray flux is negligible at 1 TeV, and the associated diffuse neutrino flux ($\sim 10^{-16}\;\mathrm{TeV\;cm^{-2}\;s^{-1}\;sr^{-1}}$) remains orders of magnitude below current IceCube limits. Nevertheless, at ultra-high energies ($>10\;\mathrm{TeV}$), this emission establishes an irreducible local background that overtakes the strongly attenuated extragalactic isotropic gamma-ray background. Our results demonstrate that the Galactic stellar ensemble is a strictly sub-dominant background, indicating that stellar subtraction templates are not required for identifying Galactic PeVatrons or constraining dark matter annihilation.

astro-ph.HE

Guardian-as-an-Advisor: Advancing Next-Generation Guardian Models for Trustworthy LLMs

Hard-gated safety checkers often over-refuse and misalign with a vendor's model spec; prevailing taxonomies also neglect robustness and honesty, yielding safer-on-paper yet less useful systems. This work introduces Guardian-as-an-Advisor (GaaA), a soft-gating pipeline where a guardian predicts a binary risk label plus a concise explanation and prepends this advice to the original query for re-inference, keeping the base model operating under its original spec. To support training and evaluation, GuardSet is constructed, a 208k+ multi-domain dataset unifying harmful and harmless cases with targeted robustness and honesty slices. GuardAdvisor is trained via SFT followed by RL to enforce label-explanation consistency. GuardAdvisor attains competitive detection accuracy while enabling the advisory workflow; when used to augment inputs, responses improve over unaugmented prompts. A latency study shows advisor inference uses below 5% of base-model compute and adds only 2-10% end-to-end overhead under realistic harmful-input rates. Overall, GaaA steers models to comply with the model spec, maintaining safety while reducing over-refusal.

cs.LG

Shorter, but Still Trustworthy? An Empirical Study of Chain-of-Thought Compression

Long chain-of-thought (Long-CoT) reasoning models have motivated a growing body of work on compressing reasoning traces to reduce inference cost, yet existing evaluations focus almost exclusively on task accuracy and token savings. Trustworthiness properties, whether acquired or reinforced through post-training, are encoded in the same parameter space that compression modifies. This means preserving accuracy does not, a priori, guarantee preserving trustworthiness. We conduct the first systematic empirical study of how CoT compression affects model trustworthiness, evaluating multiple models of different scales along three dimensions: safety, hallucination resistance, and multilingual robustness. Under controlled comparisons, we find that CoT compression frequently introduces trustworthiness regressions and that different methods exhibit markedly different degradation profiles across dimensions. To enable fair comparison across bases, we propose a normalized efficiency score for each dimension that reveals how na\"ive scalar metrics can obscure trustworthiness trade-offs. As an existence proof, we further introduce an alignment-aware DPO variant that reduces CoT length by 19.3\% on reasoning benchmarks with substantially smaller trustworthiness loss. Our findings suggest that CoT compression should be optimized not only for efficiency but also for trustworthiness, treating both as equally important design constraints.

cs.CL