arXiv ScienceSearch

arXiv subjects

Jun Fang

Publications and source records attributed to Jun Fang.

At least 19 recordsLinked to original sources

Super-Resolution Initialization of High-Fidelity CFD Simulations for Pebble-Bed Reactors

High-order CFD simulations provide detailed resolution of the heterogeneous interstitial flow in pebble-bed reactors, but their computational cost is high, especially during the initial flow-development period required to reach statistically stationary conditions. This work investigates the use of a Super-Resolution Graph Neural Network (SR-GNN) to improve the initialization of high-order NekRS simulations. Lower-order P = 2 velocity fields are used as inputs to reconstruct higher-order representations, which are then used as initial conditions for P = 7 restart simulations. The approach is evaluated using a 146-pebble bed at Re = 1000, Re = 2500, and Re = 5000, with pressure-drop convergence used as the main figure of merit. The SR-GNN models were trained using paired low- and high-order snapshots and were first evaluated through qualitative inference comparisons. High-order restart simulations showed that, for Re = 1000 and Re = 2500, the SR-GNN initialized cases produced pressure-drop histories similar to direct restarts from true P = 2 fields. For Re = 5000, however, the super-resolved field restart approached the statistically stationary P = 7 pressure-drop range faster than both the direct P = 2 restart and the reference P = 7 simulation initialized from a uniform velocity field. The trained Re = 5000 model was also applied to a larger 1568-pebble bed, demonstrating qualitative applicability of the workflow to a significantly larger packed-bed geometry. These results indicate that SR-GNN-based initialization is a promising strategy for reducing high-order flow-development cost, while also motivating further work on broader Reynolds-number and geometry generalization.

physics.flu-dyn

Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing

Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning. We therefore test whether placement and training curriculum interact. In a controlled distillation study with a Qwen3-8B teacher and a nine-layer student, pre-norm and post-norm are indistinguishable under joint training, differing by $0.0004$ validation CE, while post-norm improves over pre-norm by $0.0328$ under curriculum growth, an order of magnitude larger. A post-joint control matched by student active-layer tokens remains worse than post-grow, which rules out compute as the sole explanation. The ranking crosses over during the curriculum: post-norm takes the lead once blocks are appended. Single-block and freeze controls localize the ranking change to block appending rather than shallow-block quality or retraining. Boundary diagnostics associate post-norm with stable residual scales and pre-norm with structural-token scale drift; on a fixed batch, the final pre-grow block is also nearly identity-mapped. Together with the phase-wise crossover, these observations are consistent with boundary-scale conditioning after new blocks are appended. The results motivate treating normalization placement and training curriculum as coupled design choices in this distillation setting.

cs.AI

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization

In current zero-shot text-to-speech systems, conventional semantic tokenizers are typically optimized using supervised automatic speech recognition or self-supervised learning objectives. However, due to the inherent nature of speech, semantic and acoustic information cannot be completely decoupled, and ASR-based tokenizers discard acoustic details to focus on linguistic content; models relying on them usually struggle to achieve optimal speaker similarity. Furthermore, these tokenizers are optimized independently and lack direct supervision from downstream acoustic generation tasks. This isolated training creates a feature gap between the extracted discrete tokens and the continuous space required by acoustic models, fundamentally bottlenecking the upper bound of synthesis quality. To bridge this gap, we propose Phoenix TTS, a unified framework that tightly couples representation learning with generative acoustic modeling. Specifically, our speech tokenizer is optimized to reconstruct self-supervised features to maintain semantic richness, while simultaneously receiving direct supervision from a Flow Matching training loss. Through this joint training paradigm, the extracted discrete tokens successfully preserve essential semantic information and natively align with the feature space of the downstream Flow Matching model. Comprehensive evaluations highlight the efficiency and effectiveness of Phoenix TTS. Trained on 110K hours of data, the system achieves excellent speech intelligibility, yielding WER that consistently falls below that of ground-truth recordings. Simultaneously, it maintains robust zero-shot speaker similarity, rivaling or outperforming several prominent large-scale baselines. Furthermore, as an advantageous byproduct of this unified training, the learned tokenizer can be seamlessly adapted to zero-shot voice conversion tasks without requiring task-specific fine-tuning.

cs.SD

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.

cs.SD

Zeroth-Order Blind Interference Suppression for Multi-RIS-Aided Wireless Systems

In this paper, we study measurement-driven signal-to-interference-plus-noise ratio (SINR) maximization for a multi-reconfigurable intelligent surface (RIS)-aided single-input single-output (SISO) system with unknown strong interference sources. Specifically, the objective is to optimize the reflection coefficients such that the SINR is maximized at the receiver. As the interference channels are unknown, such an optimization problem is a black-box optimization problem with an objective function whose closed-form analytical expression is unknown. To address the high-dimensional black-box optimization problem with discrete variable constraints, we introduce a group-based phase parameterization that significantly reduces the search dimension. Building on this model, we develop a group-based zeroth-order adaptive moment (ZO-AdaMM) algorithm. Simulation results show that the proposed grouping strategy markedly accelerates the convergence speed and achieves a superior interference suppression performance under limited measurement budgets, especially in the small-budget regime.

eess.SP

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications

JD$.$com, one of the world's largest e-commerce platforms, serves over 700 million active users and millions of merchants, with a catalog of tens of billions of SKUs. At this scale, high-quality, structured item knowledge underpins a better consumer experience, lower management costs, and higher operational efficiency-yet producing and serving it poses three industrial-scale challenges: fast-emerging concepts, high-quality knowledge production for massive SKUs, and diverse downstream requirements. To address these challenges, we present the JD Oxygen AI Item Center (Oxygen AIIC), an industrial-scale platform built on LLMs/VLMs for item-knowledge production and service. Oxygen AIIC is built around four core pillars: (i) ontology engineering driven by efficient human-AI collaboration, which supports the dynamic evolution and agile expansion of an ontology with millions of entries; (ii) a "Semantic Search then Discrimination"(S2D) knowledge identification architecture that, combined with throughput improvement strategies, enables scalable, extensible, and high-throughput AI Item Library production for tens of billions of SKUs; (iii) self-evolving item-understanding LLMs/VLMs that improve in a stable and controllable manner, enabling knowledge production with 94.2% precision and 82.8% recall; and (iv) a unified item tunnel that serves as the data and service hub. Oxygen AIIC now covers tens of thousands of JD categories and processes hundreds of millions of item updates per day on Huawei Ascend NPUs. It has accumulated hundreds of billions of item-knowledge assets. Deployed across core business scenarios-including search, recommendation, operations, category planning-Oxygen AIIC has delivered measurable gains at scale. Search-traffic coverage reaches 80.4%, item-information quality issues drop by 37%, the automated fill rate of core attributes during item listing exceeds 80%.

cs.AI

BLADE: Scalable Bi-level Adaptive Data Selection for LLM Training

As Large Language Model (LLM) datasets scale to trillions of tokens, data selection has emerged as a critical frontier to filter out uninformative noise and construct adaptive learning trajectories. Beyond static heuristic filtering, advanced data selection methods for LLM training largely follow two paradigms, each with fundamental limitations. Influence-based methods provide principled bi-level objectives but require intractable inverse-Hessian computations, while excess-loss methods are computationally efficient but rely on a static reference model that becomes misaligned with the evolving proxy model during training. We propose BLADE (Bi-Level Adaptive Data sElection), a Hessian-free framework for data selection. BLADE reformulates the bi-level optimization problem underlying influence-based methods as a penalized single-level objective via Lagrange multipliers, avoiding inverse-Hessian computation while revealing a principled connection to excess-loss based data selection. The resulting objective recovers an excess-loss form but replaces the static reference model with a dynamic one that stays synchronized with training. Theoretically, we prove that this penalized formulation guarantees first-order convergence. For efficient online batch selection, we instantiate BLADE as a memoryless randomized block-coordinate Frank-Wolfe algorithm. Extensive experiments show that BLADE consistently outperforms state-of-the-art data selection baselines, providing a practical recipe for LLM training.

cs.LG

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

Zero-shot text-to-speech (TTS) relies on robust speech representations. However, current speech tokenizers face a fundamental trade-off: acoustic codecs preserve high-fidelity audio but lack linguistic constraints, causing content errors during generation, whereas semantic tokens from self-supervised learning (SSL) models ensure precise text alignment but discard some acoustic information. To bridge this gap, we propose SARA, a dual-stream VAE that directly fuses a frozen SSL semantic anchor with a dedicated residual acoustic encoder. This effectively mitigates the dilemma, creating an efficient and compact latent space without relying on complex regularizers. SARA achieves superior reconstruction quality over strong baselines. Furthermore, in downstream zero-shot TTS tasks, it yields highly natural and expressive synthesis quality, and maintains robust generation performance even under accelerated inference, offering a favorable trade-off between synthesis speed and computational cost.

cs.SD

Unlocking the Potential of Movable Antennas: General and Practical Antenna Position Optimization

Recently, movable antenna (MA) has attracted wide attention in wireless communications due to its potential in enhancing wireless communication performance via local movement within a confined region. However, antenna position optimization (APO) has emerged as a major challenge for MAs, due to the lack of a tractable, analytical, and accurate channel model in terms of antenna positions. Although existing works have developed various algorithms for APO, most of them are based on simplified theoretical channel models, which limit their generality. To address this challenge, in this article, we present more general and effective APO algorithms for different purposes, categorized as continuous APO and discrete APO, respectively. Continuous APO is mainly applied for flexible array signal processing to boost large-scale communication performance, while discrete APO is applied for small-scale multi-path channel reshaping. Specifically, the discrete APO discretizes the antenna movement region into multiple sampling points and employs discrete algorithms to determine the optimal MA positions based on the point-wise channel state information (CSI), without the need for an analytical channel model. To reduce the overhead for CSI acquisition, we also present more efficient learning-based APO algorithms that operate without requiring full point-wise CSI. Finally, we compare the application scenarios of the proposed algorithms and validate their effectiveness with numerical results.

eess.SP

Gamma-Ray Emission from the Crab Pulsar: A 17-Year Fermi-LAT Reanalysis

We present a reanalysis of 17 years of Fermi Large Area Telescope (LAT) observations of the Crab pulsar obtained between 2008 August and 2025 August. Using monthly Jodrell Bank radio ephemerides, we assigned pulse phases to the LAT events and aligned the phase zero across the full data set. From this phase-aligned data set, we derived pulse profiles over 100 MeV to 300 GeV. The pulsed emission remains clearly detectable in the 10 to 20 GeV and 20 to 30 GeV bands, with H-test significances of 32.36 sigma and 11.59 sigma, respectively, but is not significantly detected in the 30 to 300 GeV band. Phase-resolved likelihood analysis was performed over 100 MeV to 30 GeV using 14 phase bins with comparable pulsed statistics. The fixed-window fractional fluxes show that the contribution of Peak 1 (P1) decreases steadily with energy, while those of Peak 2 (P2) and the Bridge increase, with P2 exceeding P1 above 10 GeV. Finally, the same phase-assignment framework also enables an off-pulse analysis from 100 MeV to 1 TeV, confirming the synchrotron and inverse-Compton components that dominate the emission in the selected off-pulse interval.

astro-ph.HE

Mixture-of-Experts Diffusion Models for Adaptive Massive MIMO Channel Estimation via Variational Bayesian Inference

Channel estimation is essential to massive multiple-input multiple-output (MIMO) systems. While recent generative model-based approaches using lightweight diffusion models (DMs) have achieved superior performance, they typically rely on a single data-driven prior, which limits their adaptability to varying channel distributions in real-world scenarios. To address this deficiency, we propose a mixture-of-experts (MoE) diffusion model (DM) framework combined with variational Bayesian inference. Specifically, our approach employs multiple pre-trained DMs, with each trained on a specific type of propagation channels. We then propose a probabilistic graphical model in which the channel is modeled as a latent variable drawn from one of these candidate generative priors with a certain probability. By integrating variational Bayesian inference with DM-based data priors, the underlying channel along with the expert indicator variable are jointly inferred, thus enabling automatic model adaptation for channel estimation. The effectiveness of our approach is evaluated on 3GPP CDL channels. Simulation results demonstrate that our proposed approach achieves a clear performance improvement over the standard DM-based method that employs a single prior trained on aggregated data from all channel types, particularly when the channel samples from different propagation environments are imbalanced.

eess.SP

Forward Modeling of Dust-Induced Stray Light in Ground-Based Coronagraphs: A Dual-Path Monitoring Approach for High-Precision Inner Corona Observations

High-precision ground-based observations of the inner corona (1.05-2.0 R_sun) are fundamentally constrained by instrumental stray light, particularly the additive background from dynamic dust accumulation on the objective lens. To address this issue, we propose a correction method for the Spectral Imaging Coronagraph (SICG) based on dual-path real-time monitoring and forward physical modeling. By simultaneously imaging the objective lens surface, we obtain deterministic prior information on dust distribution. We construct a physical point-spread function using optical defocus parameters and reconstruct the nonuniform scattering background via convolution. Model parameters are retrieved through data-driven inversion constrained by polar coronal holes. The method demonstrates excellent robustness under varying contamination conditions. After correction, the rms noise in the polar background is reduced by approximately 67% on average, and the signal-to-background ratio improves by a factor of up to 3.7 under heavy contamination conditions. Comparisons with space-based Solar Dynamics Observatory/Atmospheric Imaging Assembly observations indicate that the corrected images recover the morphological structures of streamers with high fidelity. Further radial intensity analysis reveals that the correction process successfully restores the hydrostatic exponential decay characteristic of inner coronal radiation. The fitted decay coefficient corresponds to a plasma temperature of approximately 2.0 MK, consistent with the characteristic formation temperature of the Fe XIV 530.3 nm line. These results demonstrate that the method effectively eliminates the dominant systematic bias in ground-based observations, providing a reliable data foundation for high-precision coronal thermodynamic and dynamic research with the SICG.

astro-ph.SR

Cassie-Wenzel transition induced by localized freezing after droplet impact on supercooled micro-patterned surfaces

Micro-patterned surfaces have attracted significant attention in numerous applications owing to their potential to enhance hydrophobic and icephobic properties. A Cassie state of final wetting of a droplet upon impact on a micro-patterned surface, which is highly favorable for anti-icing applications, is achieved in this study through rapid localized freezing in the droplet-surface contact region via tuning the coupled interplay among droplet spreading kinetics, interfacial heat transfer, and solidification dynamics. Synchronized high-speed imaging and infrared thermography are employed to probe droplet impact and freezing dynamics, with particular emphasis on the transition of wetting state and its effect on the resulting freezing morphology. Experimental results reveal that variations in impact velocity and wall temperature lead to a final frozen wetting-state transition of the droplet from the Wenzel to the Cassie regime, accompanied by pronounced changes in freezing time, final spreading diameter, and frozen height. The transition of wetting states is attributed to rapid localized freezing at the droplet bottom, which suppresses liquid penetration into the micro-pattern. At lower impact velocities and surface temperatures, droplets tend to maintain the Cassie state with extended freezing durations, whereas higher velocities or higher temperatures promote rapid penetration and accelerated freezing. This study elucidates the coupled penetrating-freezing mechanism governed by micro-pattern design and provides fundamental insights into the rational design of anti-icing and icephobic surfaces.

physics.flu-dyn

StoryEcho: A Generative Child-as-Actor Storytelling System for Picky-Eating Intervention

Picky eating in children can undermine dietary diversity and the development of healthy eating habits, while also creating recurring tension in family feeding routines. Prior interventions have explored food-centered designs, enhanced utensils, and mealtime interactive systems, but few position children as active participants in intervention processes that extend beyond single mealtime interactions. To better understand everyday responses to picky eating and child-acceptable intervention mechanisms, we conducted a formative study with caregivers and kindergarten teachers. Based on the resulting design considerations and iterative stakeholder review, we designed StoryEcho, a generative child-as-actor storytelling system for picky eating intervention. StoryEcho engages children outside mealtimes through personalized stories in which the child appears as a persistent story character and later shapes story development through real-world food-related behavior. The system combines non-mealtime story engagement, lightweight post-meal feedback, and behavior-informed story updates to support repeated intervention across everyday family routines. We evaluated StoryEcho in a between-group field study with 11 families of preschool children. Results provide preliminary evidence that StoryEcho can significantly increase children's willingness to approach and try target low-preference foods while reducing parental pressure around feeding. These findings suggest the promise of generative child-as-actor storytelling as a design approach for home-based behavior support that unfolds through recurring family routines.

cs.HC

SpeakSoftly: Scaffolding Nonviolent Communication in Intimate Relationships through LLM-Powered Just-In-Time Interventions

Conflicts are common in text-based communication, particularly in intimate relationships, where misunderstandings can easily escalate into verbal aggression. To address this, we present SpeakSoftly, a system that applies Nonviolent Communication (NVC) principles to scaffold couples' conflict communication through LLM-powered just-in-time interventions. Informed by formative interviews with couples and NVC principles, we designed two core features: NVC-Prompt, which detects verbal aggression and suggests revisions to prevent escalation, and NVC-Guide, which analyzes dialogues to uncover users' feelings and needs, fostering self-awareness and perspective-taking. These features were implemented across three progressive intervention modes, each varying in intervention depth and tone: Basic Reminder, Neutral Guide, and Empathetic Guide. We conducted a mixed-methods user study with 18 couples across simulated and real-life conflict settings to evaluate the effectiveness of each mode. Results showed that Empathetic Guide significantly facilitated both behavioral and cognitive changes, while Neutral Guide was effective only for behavioral changes in simulated conflicts. In real-life conflicts, Neutral Guide showed distinct advantages due to lower cognitive load demands. We discuss the mechanisms behind these findings and propose design implications for in-situ interventions in high-stakes communication contexts.

cs.HC

Efficient Solving for Dynamic Data Structure Constraint Satisfaction Problem

Functional verification plays a central role in ensuring the correctness of modern integrated circuit designs, where constrained-random verification is widely adopted to generate diverse stimuli under high-level constraints. In industrial verification environments, constraint solving increasingly involves dynamic data structures whose shape and content are determined at runtime, causing the sets of variables and constraint instances to evolve across solver invocations, which in turn leads to substantial overhead when nested and high-dimensional structures repeatedly expand across solves. We formalize this class of problems as the Dynamic Data Structure Constraint Satisfaction Problem (D2SCSP),which captures the interaction between dynamic data structure expansion and constraint evaluation. We propose a dependency-guided problem partitioning framework combined with an incremental encoding and constraint activation mechanism, enabling reuse of solver state and encodings across multiple solves. The framework is integrated into an industrial SystemVerilog verification flow and implemented in the commercial simulator VeriSim. Experimental results on industrial benchmarks demonstrate significant performance improvements, achieving an average speedup of 24.80x over a baseline and 1.72x over a state-of-the-art commercial simulator, highlighting the practicality of the approach for real-world verification workflows.

cs.AR

Routine Computing: A Systematic Review of Sensing Daily Life Dimensions Towards Human-Centered Goals

Human routines structure daily life, yet remain challenging for computational systems to understand. This paper presents the first systematic review of routine computing, a previously implicit but increasingly recognized field that focuses on computationally sensing and modeling human behaviors. It synthesizes 203 studies published up to August 2025. The paper presents a new taxonomy of the literature, focusing on temporal structures, behavioral interactions, cognitive aspects, and how variability and deviations are addressed. The common goals of routine computing extend across four major application domains, including accessibility care, the promotion of healthy habits, adaptive and context-aware support, and large-scale population insights. Persistent challenges that limit the design of truly human-centered systems are identified, including the gap between low-level activity recognition and high-level intent, the tension between personalization and generalization, unresolved privacy concerns, and data-related limitations. By consolidating these findings, this paper provides a foundational framework for HCI researchers, outlining principles for designing ethical, adaptive, and human-centered routine-aware systems.

cs.HC

Terahertz Beam Squint Mitigation via Six-Dimensional Movable Antennas

Analog beamforming holds great potential for future terahertz (THz) communications due to its ability to generate high-gain directional beams with low-cost phase shifters. However, conventional analog beamforming may suffer substantial performance degradation in wideband systems due to the beam squint effect. Instead of relying on high-cost true-time delayers, we propose an efficient six-dimensional movable antenna (6DMA) architecture to mitigate the beam-squint effect. In particular, we study a wideband wide-beam coverage problem in this paper, aiming to maximize the minimum beamforming gain over a given range of azimuth/elevation angles and frequencies by jointly optimizing the analog beamforming vector, the MA positions within a two-dimensional (2D) region, and the three-dimensional (3D) rotation angles of the antenna array. However, this problem is non-convex and intractable to solve optimally due to the coupling of the spatial and frequency domains and that of the antenna weights, positions and rotation. To tackle this problem, we first derive an optimal solution to it in a special case with azimuth or elevation angle coverage only. It is shown that rotating a uniform linear array (ULA) is sufficient to achieve global optimality and eliminate beam-squint effects. While for other general cases, an alternating optimization (AO) algorithm is proposed to obtain a high-quality suboptimal solution, where the antennas' beamforming weights, positions, and rotation angles are alternately optimized by combining successive convex approximation (SCA), sequential update with Gibbs sampling (GS), and hybrid coarse- and fine-grained search. Simulation results demonstrate that our proposed scheme can significantly outperform conventional antenna arrays without antenna movement or rotation, thus offering a cost-effective solution for wideband transmission over THz bands.

eess.SP