arXiv ScienceSearch

arXiv subjects

Yan Sun

Publications and source records attributed to Yan Sun.

At least 19 recordsLinked to original sources

Star Formation in the H II Region Sh 2-205: 3D Morphology and Kinematics from Young Stars and Molecular Gas

Using Gaia astrometry of young stars combined with CO observations, we present the first systematic three-dimensional (3D) analysis of the structure, kinematics, and evolutionary history of the star-forming regions in the environs of the H II region Sh 2-205 (S205). S205 exhibits a complex morphology and coherent expansion on both global and subregional scales. We identify several O9-B1 stars and a 0.56 Myr old pulsar that are likely associated with the region. A momentum estimate suggests that feedback from these objects may account for the observed overall expansion. Trace-back analysis of the expansion, combined with color-magnitude diagram fitting for young star clusters, indicates at least two episodes of star formation. These results reveal a complex star-formation history of S205 and provide new insights into its 3D evolution.

astro-ph.GA

UniDot: A Unified Network for Sequence Modeling and Feature Interaction in Large-scale Recommendation

Industrial recommenders rely on two model families that have evolved largely independently: feature-interaction models over multi-field user/item features, and sequential models over user-behavior histories. Production systems couple them only loosely. To unify the two, we present UniDot, a novel architecture for post-click conversion prediction built from the factorization-machine (FM) point of view: the embedding inner product---which powers collaborative filtering and lets a recommender generalize to unseen user--item pairs---is the same primitive as attention's query dot key scoring, so a single dot-product of tokens can underlie both feature interaction and sequence modeling. UniDot tokenizes non-sequential fields and multi-domain behavioral sequences into one shared token space and stacks a single macro-block in which a token-mixing bus and a sequence-retrieval bus (item tokens cross-attending the histories) run in parallel and exchange state each layer through an MLP-Mixer fusion, while an FM Highway carries explicit per-layer dot-product interactions around the residual stack directly to the classifier. The sequence side is embedded once per forward pass and shared by all consumers, bounding inference latency. Trained with a dual sparse/dense (Adagrad + Muon) optimizer, an auxiliary conversion-delay head, and multi-path mutual learning, UniDot finished as the runner-up on the Industrial track of the TAAC KDD Cup 2026.

cs.IR

Trace the Self-Gravitating Gas Using CO Isotopologues

Recent studies have shown that the star formation rate (SFR) correlates tightly and linearly with the mass of gravitationally bound gas, which can be delineated from the power-law tail of the column-density probability distribution function ($N$-PDF) derived from dust emission observations. This relationship holds across four orders of magnitude within the Milky Way--spanning low-mass to high-mass star-forming regions and encompassing the extreme environment of the Central Molecular Zone. Building on this framework, we present a new approach for estimating the mass of gravitationally bound gas in molecular clouds using multi-line CO isotopologue observations. Our sample includes 16 molecular clouds with robust detections in $^{12}$CO, $^{13}$CO, and C$^{18}$O $J$ = 1-0, spanning both massive inner Galaxy clouds and nearby star-forming regions. We find that the $N$-PDFs derived from combined CO isotopologue data recover the characteristic log-normal plus power-law profiles seen in dust-based studies. The mass and spatial distribution of the self-gravitating structures estimated from both dust-based and CO-based methods agree well throughout the sample. This indicates that the CO isotopologue combination can robustly trace the self-gravitating component via the $N$-PDF method and provides a reliable, scalable, and velocity-resolved alternative to dust emission for identifying the star-forming gas in molecular clouds.

astro-ph.GA

CO Structures with Narrow Lines in Nearby Quiescent Regions

Using CO data from Phase I of the Milky Way Imaging Scroll Painting (MWISP) survey, we present a systematic study of molecular structures with narrow lines. We identify 57 CO structures, most of which exhibit low densities and subsonic/transonic turbulence. Among them, structures with large projected areas and diffuse, sheet-like geometries are identified as veil clouds. The low LSR velocities and the concentration of these CO structures toward both the Galactic center (e.g., Ophiuchus, Aquila) and anticenter (e.g., Cepheus, Taurus) regions suggest a local origin for the sample, as supported by distance measurements of about 200--300pc for a subset with relatively large angular extents. These nearby structures likely arise from large-scale compression driven by past supernova activity within the Local Bubble. The observed low-velocity-dispersion emission may trace quiescent regions where turbulence has decayed due to a lack of sustained energy injection. For diffuse veil clouds with an assumed magnetic field of ~10uG, ion-neutral friction may provide an additional mechanism for turbulent dissipation on sub-parsec scales corresponding to their thickness of 0.1--0.3pc. Tracing the atomic-to-molecular transition, veil clouds provide a unique window into the diffuse, quiescent precursor state of dense gas. They likely represent a widespread but previously overlooked component of the Galactic molecular gas reservoir, with significant implications for cloud formation and evolution, the total mass budget and spatial distribution of molecular gas, and the initial conditions of star formation as a related consequence.

astro-ph.GA

Investigations of MWISP Bubbles: Identification and Analysis of Enclosed Molecular Bubbles by Weight Fields

Molecular bubbles are widely used as tracers of stellar feedback; yet, their identification in spectral-line surveys remains challenging because both cavity morphology and kinematic structure must be assessed consistently in position--position--velocity (PPV) space. We present the Bubble-Weight Fields (BWFields) framework, a PPV-based method that for the first time enables the automated and objective identification and analysis of enclosed molecular bubbles directly from spectral-line data cubes. BWFields constructs a bubble-weight field, $W_{l,b,v}$, which encodes cumulative evidence for cavity interiors by aggregating topological signatures across multiple signal-to-noise tiers and velocity-integration scales. Contiguous cavity interiors are segmented as weight-clumps and associated with surrounding molecular gas, linking candidate bubbles to the structure of their host clouds. Shell morphology is characterized using radial intensity profiles and emission-defined intensity skeletons, which capture the shell geometry as traced by the observed emission. Bubble kinematics are quantified using azimuthally sampled position-velocity (PV) diagnostics, along with a turbulence-normalized expansion significance, which serves as a direct measure of the expansion-like velocity organisation. Applied to MWISP $^{13}$CO observations of the G17 region, BWFields identifies a population of bubble candidates with a broad range of morphologies and velocity structures in complex environments. BWFields establishes a scalable and physically interpretable framework for molecular-bubble studies in large surveys, enabling systematic investigations of stellar feedback in the Galactic interstellar medium.

astro-ph.GA

TFGformer: Multivariate Time Series Forecasting via Time-Frequency Graph Learning and Covariate Fusion

Large-scale multivariate time series from heterogeneous IoT sensors demand accurate long-term forecasting for resource scheduling and predictive maintenance. While recent time series foundation models exhibit strong generalization, they rely on static parametric knowledge and lack dynamic access to external historical patterns during inference. Retrieval-Augmented Generation (RAG) offers a potential remedy, yet its application to time series forecasting is challenged by magnitude variations across heterogeneous sources and the mismatch between historical similarity and future consistency. We propose CrossRAG, a retrieval-augmented forecasting framework that integrates Shape-Aware Memory (SAM) with RevIN normalization for magnitude-robust shape-level retrieval, Future-Consistent Contrastive (FCC) learning to distinguish informative references from hard negatives with similar history but divergent futures, and Cross-Attention Temporal Fusion (CATF) to fuse retrieved historical--future reference pairs into the backbone's representations at the representation level. Experiments on seven public benchmarks show that CrossRAG consistently outperforms both parametric-only baselines and existing retrieval-augmented forecasting methods.

cs.LG

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state. Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift toward agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time. To address this gap, we introduce JarvisHub, a canvas-native creative agent harness for long-horizon multimodal creation. JarvisHub treats an editable canvas as the user workspace, the agent's external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use toward sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.

cs.CV

Ai2-Kit: Streamlining AI-Accelerated Ab Initio Workflows for Complex Chemical Systems

Molecular simulations of complex chemical systems, such as catalysis, electrochemistry, and energy storage, often need to capture the interplay of effects such as electronic structure, finite-temperature fluctuations, and electric-field response. Such complexity is difficult to address with traditional ab initio calculations, which are limited by the time and length scales they can reach. AI-accelerated ab initio (AI2) methods use machine learning potentials trained on first-principles data to replace expensive electronic-structure calculations, extending ab initio accuracy to these regimes, but their routine application requires reliable workflows that connect first-principles calculations, model training, molecular dynamics, enhanced sampling, trajectory analysis, and HPC orchestration. Here we present ai2-kit, a software toolkit for developing accessible, reproducible, and extensible AI2 workflows. ai2-kit provides high-semantic-density command-line interfaces and Python APIs for structure and dataset conversion, batch task generation, active-learning screening, job orchestration, and workflow recovery. We demonstrate ai2-kit in four representative applications: active-learning-based machine learning potential construction, free-energy perturbation for redox and acid-base processes, electrochemical machine learning potentials for electrified interfaces, and spectroscopies from machine learning molecular dynamics. ai2-kit also provides AI-agent skills that help users adapt these use cases into customized workflows for their own chemical systems and computational software stacks. Together, ai2-kit helps turn AI2 methods from bespoke computational protocols into reusable and extensible workflows for complex chemical systems, from model construction to property prediction.

physics.chem-ph

Statistical Properties of Molecular Clouds in the Milky Way: Insights from Three-Isotopologue CO Observations of the MWISP Project

We present a comprehensive statistical analysis of molecular cloud (MC) properties using the MWISP survey's 12CO, 13CO, and C18O (J = 1--0) data toward the inner (l = 45$^\circ$--60$^\circ$) and outer (l = 120$^\circ$--130$^\circ$) Galaxy. From a strict selection of 24,724 identified MCs, a final sample of 3,161 well-resolved MCs is established. We investigate the distributions of observational, morphological, and derived physical parameters, as well as their environmental dependencies and intercorrelations. Our analysis reveals that MCs are typically oblate and tend to align with the Galactic disk. A critical evaluation using a nearby subsample confirms significant distance-dependent selection effects for some parameters, nevertheless, the direction of changes in these parameters can indicate distance influence. We also examine several specific subsamples, revealing the distinct characteristics of MCs in the G120 spiral shock region, MCs in the G50 interarm spurs, C18O-bright MCs, and MCs with supra-Larson velocity dispersion. For instance, MCs with supra-Larson velocity dispersion are predominantly small and likely young clouds inheriting turbulence from the diffuse ISM. Notably, a comparison across tracers reveals that typical MCs have a turbulent, diffuse, 12CO-bright gas structure in their outer layers that does not contribute directly to star formation. In contrast, 13CO-bright gas represents a turning point where gravity becomes significant; C18O-bright gas is about gravity-dominated. Comprehensive correlation analysis confirms a flatter $\sigma_v$-size relation than classic Larson's law and a strong mass-size relation. Incorporating dimensional analysis, we derive minimal sets of eigenparameters from which most other observational and physical parameters can be estimated. This highlights the underlying scaling relations that governing cloud properties.

astro-ph.GA

Sublinearly Structured Deep Neural Networks Achieve Feature Learning Consistency for Compositional Functions

Over the past decade, deep neural networks (DNNs) have achieved remarkable success on complex machine-learning tasks, yet the theoretical foundations of their performance remain incomplete. From a statistical viewpoint, a natural question is: can DNNs attain feature-learning and prediction consistency comparable to that of classical models? While a full characterization is open, we provide positive results for a broad subclass. We establish feature-learning consistency guarantees for sublinearly structured DNNs-architectures whose input/output dimensions and number of hidden neurons grow sublinearly with the sample size-when learning hierarchically compositional target functions. Importantly, this consistency still holds even in the conventional "over-parameterized" regime where the total number of parameters exceeds the number of training samples. Empirically, sublinearly structured DNNs match or surpass wide DNNs in prediction. A structural audit further indicates that widely used convolutional neural networks (CNNs), including AlexNet, VGGNet, ResNet, GoogLeNet, are sublinearly structured on their image classification benchmarks. We further prove that the sublinearly structured DNNs achieve universal approximation for hierarchically compositional functions in the large-sample limit. Moreover, images exhibit an inherent hierarchical, compositional structure. Taken together, these results explain, through a statistical lens, why many large-scale deep learning models succeed after adequate training on massive image datasets.

stat.ML

OneReason Technical Report

Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce. However, these generative models can only benefit from the scaling advantage, while their reasoning ability is hard to activate, since we cannot construct meaningful Chain-of-Thought (CoT) sequences consisting of itemic tokens only. Inspired by the success of the reasoning-style ``think before answer'' paradigm in the LLM field, we conduct preliminary studies (i.e., OneRec-Think, OpenOneRec) to explore reasoning capability in generative recommendation. Nevertheless, we notice an unexpected phenomenon: the thinking mode does not show advantages over the non-thinking mode. Drawing insights from recent findings on CoT robustness in multi-modal language models, we argue that effective reasoning in recommendation rests on two factors: perception, the ability to ground itemic tokens in their underlying language semantics, and cognition, the ability to reorganize a user's behavior sequence into coherent latent interest points. We therefore propose OneReason, which includes: (1) strong itemic token perception in pre-training, (2) a three-level cognition-enhanced CoT format for recommendation tasks in SFT, and (3) a specialize-then-unify training recipe in RL to enhance the thinking ability.

cs.IR

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization

Pretraining large language models (LLMs) with next-token prediction has led to remarkable advances, yet the context-dependent nature of token embeddings in such models results in high intra-class variance and inter-class similarity, thus hindering the efficiency of representation learning. While similarity-based regularization has demonstrated benefit in supervised fine-tuning and classification tasks, its application and efficacy in large-scale LLM pretraining remains underexplored. In this work, we propose the SimReg, an embedding similarity regularization loss that explicitly encourages token representations with the same ground-truth label within each sequence to be more similar, while enforcing separation from different-label tokens via a contrastive loss. Our analysis reveals that this mechanism introduces gains by enlarging multi-classification margins, thereby enabling more efficient classification. Extensive experiments across dense and Mixture-of-Experts (MoE) architectures demonstrate that SimReg consistently accelerates training convergence by over 30% and improves average zero-shot downstream performance by over 1% across standard benchmarks. Further ablation studies and analyses offer practical insights into hyperparameter tuning and loss effectiveness.

cs.CL

Where to Spend Rollouts: Hit-Utility Optimal Rollout Allocation for Group-Based RLVR

Reinforcement learning with verifiable rewards (RLVR) has emerged as a central paradigm for improving the reasoning capabilities of large language models. Group-based policy optimization methods, such as GRPO, typically allocate a fixed number of rollouts to every prompt. This uniform allocation can be inefficient: it over-allocates compute to prompts whose sampled groups are already saturated while under-exploring prompts for which additional samples may reveal useful correct trajectories. To address this limitation, we introduce hit utility, the posterior probability that at least one rollout in a proposed additional allocation for a prompt will be correct. Building on this notion, we propose Hit-Utility Optimal Rollout Allocation (HORA), a learning-free rollout allocation policy that maximizes total posterior hit utility within each allocation batch. HORA adaptively reallocates rollout budgets while leaving the downstream reward evaluation and group-based advantage estimator unchanged. Across four mathematical reasoning benchmarks and three model scales, HORA preserves comparable Pass@1 and improves Pass@K over compute-matched GRPO in ten of twelve model--benchmark configurations, with one tie and one saturated exception. It is also drop-in compatible with other group-based estimators such as RLOO. Ablation studies indicate that the uniform prior used by HORA is competitive with five prompt-conditioned learned-prior alternatives.

cs.LG

East Asian VLBI Network astrometry toward the star-forming region G040.96+02.48 in the Extreme Outer Galaxy

Accurate astrometric measurements for star-forming regions located on the far side of the Milky Way remain scarce. In this work, we present the astrometric results for a 22\,GHz water maser associated with star-forming region G040.96+02.48 located on the far side of the Milky Way, using the East Asian VLBI Network. The target water maser's proper motion was determined to be ($\mu_{\alpha}\cos\delta, \mu_{\delta}$) = ($-2.06_{-0.51}^{+0.53}$, $-2.95_{-0.44}^{+0.45}$)~mas~yr$^{-1}$. The derived three-dimensional kinematic distance to the star-forming region is 20.2$\pm$3.2\,kpc, placing it slightly outside the Outer Scutum$-$Centaurus Arm. The corresponding vertical height of 872$\pm$139\,pc indicates a significant warp of the outer Galactic disk, which is in good agreement with the latest precessing warp model. Moreover, the resulting peculiar motions reveal a complex kinematic pattern, characterized by a large outward radial velocity of $-32\pm$18\,km~s$^{-1}$. Our observations substantially expand the valuable sample of star-forming regions with accurate astrometric measurements in the Extreme Outer Galaxy.

astro-ph.GA

Bayesian Analysis of Gravitational Wave Microlensing Effects from Galactic Double White Dwarfs

Gravitational waves (GWs) from the galactic double white dwarf (DWD) systems are one of the primary targets for upcoming space-based detectors. Due to their vast abundance and widespread distribution throughout the Galactic disk and bulge, these systems may provide a high-statistical population for probing GW microlensing effects induced by Galactic compact objects. To evaluate the detectability of such effects, in this work we simulate the four-year observation of DWD systems by Taiji, in the form of a second-generation Time Delay Interferometry (TDI) data stream. Within a Bayesian inference framework, we estimate parameters for lensed GWs from DWD systems for different values of the lens parameters, including the lens mass $M_\mathrm{L}\in [10, 10^6]$\,M$_\odot$, the effective velocity $v_\mathrm{eff}\in [50, 500]$\,km/s and the initial separation $L\in [R_\mathrm{E}, 3R_\mathrm{E}]$, and obtain the uncertainties of the corresponding parameters. These results characterize the capability of future Taiji observations to probe such systems. We further employ the Bayesian model selection framework to distinguish between lensed and unlensed scenarios, and investigate the impacts of three key physical parameters of the lens system: $M_\mathrm{L}$, $v_\mathrm{eff}$, and $L$ on distinguishing lensing events. Our results show that when $M_\mathrm{L}$ is below $10^5$\,M$_\odot$ or $L\geq3R_\mathrm{E}$, it is not possible to distinguish between lensed and unlensed models. For $v_\mathrm{eff}$, although the Bayes factor decreases as $v_\mathrm{eff}$ decreases, the lensed and unlensed models can still be distinguished within our parameter range.

astro-ph.GA

Rethinking the Personalized Relaxed Initialization in the Federated Learning: Consistency and Generalization

Federated learning (FL) is a distributed paradigm that coordinates massive local clients to collaboratively train a global model via stage-wise local training processes on the heterogeneous dataset. Previous works have implicitly studied that FL suffers from the ``client-drift'' problem, which is caused by the inconsistent optimum across local clients. However, till now it still lacks solid theoretical analysis to explain the impact of this local inconsistency. To alleviate the negative impact of ``client drift'' and explore its substance in FL, in this paper, we first propose an efficient FL algorithm FedInit, which allows employing the personalized relaxed initialization state at the beginning of each local training stage. Specifically, FedInit initializes the local state by moving away from the current global state towards the reverse direction of the latest local state. Moreover, to further understand how inconsistency disrupts performance in FL, we introduce the excess risk analysis and study the divergence term to investigate the test error in FL. Our studies show that optimization error is not sensitive to this local inconsistency, while it mainly affects the generalization error bound. Extensive experiments are conducted to validate its efficiency. The proposed FedInit method could achieve comparable results compared to several advanced benchmarks without any additional training or communication costs. Meanwhile, the stage-wise personalized relaxed initialization could also be incorporated into several current advanced algorithms to achieve higher generalization performance in the FL paradigm.

cs.LG

A Comparative Study of TeV Gamma-Ray Sources with Various Objects

We investigate the relationships between LHAASO TeV gamma-ray sources and various kinds of objects, including pulsar wind nebulae (PWNe), supernova remnants (SNRs), HII regions, microquasars, and OB associations. We propose a Randomization-Adjusted Overlap Correlation (RAOC) method to statistically assess association probabilities and evaluate association proportions across catalogs. The results reveal statistically significant overlaps between LHAASO sources and SNRs, PWNe, and microquasars, supporting their role as important contributors to TeV gamma-ray emission. The estimated association proportions of LHAASO sources are 0.19$\pm$0.08 with SNRs, 0.20$\pm$0.04 with PWNe, and 0.027$\pm$0.008 with microquasars. The proportion of the gamma-ray sources associated with the subsample of shell-type SNRs is ~0.1. While HII regions also show potential association, particularly with the KM2A component, their large self-overlap ratio complicates precise estimation. In contrast, OB associations exhibit a high probability of chance coincidence, suggesting their limited contribution to TeV gamma-ray emission. Our analysis of TeV gamma-ray emission capabilities shows that ~60% of PWNe are gamma-ray bright in both the WCDA and KM2A energy ranges. For SNRs and microquasars, the TeV gamma-ray bright fraction is ~10%. The subsample of PWNe associated with molecular clouds (MCs) shows enhanced gamma-ray emission. Furthermore, positional analysis reveals a systematic offset of the gamma-ray sources overlapping with PWNe toward the associated MCs. These findings imply a role for MCs in PWN gamma-ray production. Additionally, self-correlation analysis indicates that about 70% of the WCDA and KM2A gamma-ray components share a common origin. The study also identifies selection effects in existing SNR catalogs and notes clustering among approximately 30% of HII regions within larger star-forming regions.

astro-ph.HE

VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection

Time series anomaly detection (TSAD) is essential for maintaining the reliability and security of IoT-enabled service systems. Existing methods require training one specific model for each dataset, which exhibits limited generalization capability across different target datasets, hindering anomaly detection performance in various scenarios with scarce training data. To address this limitation, foundation models have emerged as a promising direction. However, existing approaches either repurpose large language models (LLMs) or construct largescale time series datasets to develop general anomaly detection foundation models, and still face challenges caused by severe cross-modal gaps or in-domain heterogeneity. In this paper, we investigate the applicability of large-scale vision models to TSAD. Specifically, we adapt a visual Masked Autoencoder (MAE) pretrained on ImageNet to the TSAD task. However, directly transferring MAE to TSAD introduces two key challenges: overgeneralization and limited local perception. To address these challenges, we propose VAN-AD, a novel MAE-based framework for TSAD. To alleviate the over-generalization issue, we design an Adaptive Distribution Mapping Module (ADMM), which maps the reconstruction results before and after MAE into a unified statistical space to amplify discrepancies caused by abnormal patterns. To overcome the limitation of local perception, we further develop a Normalizing Flow Module (NFM), which combines MAE with normalizing flow to estimate the probability density of the current window under the global distribution. Extensive experiments on nine real-world datasets demonstrate that VAN-AD consistently outperforms existing state-of-the-art methods across multiple evaluation metrics.We make our code and datasets available at https://github.com/PenyChen/VAN-AD.

cs.LG