arXiv ScienceSearch

arXiv subjects

Guodong Li

Publications and source records attributed to Guodong Li.

At least 19 recordsLinked to original sources

ALOHA IRDCs Molecular Line Follow-up: I. Gas properties and kinematics

Infrared Dark Clouds are ideal sites for investigating the initial conditions of massive star and cluster formation. The A Lei Of the Habitat and Assembly of Infrared Dark Clouds (ALOHA IRDCs), a James Clerk Maxwell Telescope (JCMT) Large Program, has mapped nearby IRDCs with SCUBA-2. Complementary molecular line observations are needed to characterise the physical, kinematic, and chemical properties of the dense gas. We aim to determine the thermal, kinematic, and chemical properties of clumps identified in the ALOHA IRDCs, and to assess their evolutionary status and level of star-forming activity. We performed single-pointing K-band and W-band observations towards 56 ALOHA IRDCs clumps using the Effelsberg 100-m and Yebes 40-m telescopes, respectively. We derived NH3 kinetic temperatures using the hyperfine group ratio (HFGR) method and identified infall and shock signatures from HCO+, H13CO+, SiO, and HNCO profiles. Water masers and NH2D emission were used as complementary tracers of chemical evolution and star formation. The clumps exhibit kinetic temperatures of 15-29 K. We detect NH2D emission towards 18 sources, with NH2D centroid velocities consistent with NH3, indicating both species trace the same dense gas component. More than half of the clumps display blue-asymmetric HCO+ profiles, identifying them as infall candidates. Water masers are detected in 22 sources, with prominent velocity ranges and variability. Broad SiO emission (>~20 km/s) indicates strong shocks, while narrower extents (<~6km/s) likely trace large-scale interactions or low-velocity shocks. The widespread infall signatures, shock tracers, masers, and NH2D emission suggest that relatively quiescent, chemically young material can coexist with dynamically active gas affected by early protostellar feedback, providing insight into the coupled physical and chemical evolution of massive IRDC clumps.

astro-ph.GA

Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio cues needed to localize and track individual sources. To evaluate this missing capability, we introduce ST-OmniQA, a spatio-temporal audio-visual question-answering benchmark built from panoramic videos paired with synchronized first-order Ambisonics (FOA) audio of moving sound sources. It contains 40K videos and 400K question-answer pairs organized into four capability levels covering sound-event recognition, direction of arrival, source distance, motion trajectories, and temporally grounded audio-visual reasoning. Building on this benchmark, we propose ST-Omni-R1, which integrates FOA-derived semantic and trajectory representations with panoramic visual context and is trained through progressive curriculum learning and reasoning-tree reinforcement learning. ST-Omni-R1 achieves 77.83\% average semantic accuracy across the four levels, compared with 37.28\% for the best evaluated baseline. Results on three public spatial-audio benchmarks further indicate that its learned spatial and motion representations transfer beyond ST-OmniQA.

cs.AI

Mixed-Frequency Time Series Forecasting via Depth-Separable Neural Networks

To better forecast mixed-frequency time series, it is the key to choose a suitable way for frequency alignment. However, the existing methods are all limited to linear transformations, and this may overlook the possible nonlinearity, leading to a worse prediction. We alternatively consider a deep neural network for each frequency alignment, and hence a depth-separable neural network. Moreover, a parameter-sharing mechanism is adopted across the alignment at each stage, making possible a deeper network for a large set of higher-frequency predictors. This paper establishes an approximation theory for the proposed depth-separable network, and a non-asymptotic prediction error bound is also derived. Simulation studies demonstrate the finite-sample performance of the proposed method, and an empirical application to forecasting U.S. quarterly macroeconomic variables using monthly and daily indicators, highlights its superior predictive accuracy over existing mixed-frequency methods.

stat.ME

A Deep Study of the Spiral Galaxy W2246f

In this era of large surveys and statistical studies of galaxies, the beauty in the details of individual galaxies is often lost. We present a deep study of the spiral galaxy W2246f with MUSE, exploring the spatially-resolved stellar and ionized gas properties to understand how it formed and evolved over time. The unusually deep observations of this galaxy give us a rare opportunity to study this phenomenon with better spatial resolution than can normally be achieved with the current IFU surveys of galaxies at a similar redshift ($z\sim0.09$). We analyse the stellar and gas kinematics, as well as the spatially resolved stellar populations and gas properties, including gas metallicity and the dominant ionization sources. The derived properties include the stellar mass, radial profiles of luminosity- and mass-weighted mean ages and metallicities, and ionized gas characteristics such as E(B-V), H$\alpha$ extinction, dust-corrected H$\alpha$ flux, oxygen abundance using the O3N2 calibrator, and H$\alpha$-based star formation rate. Analysis of the stellar populations revealed a negative metallicity gradient, and the mass-weighted ages showed uniformly flat ages across the disc while the luminosity-weighted ages show a negative gradient. We find that the gas metallicity and star formation rate density also drop in the central region of the galaxy where the older luminosity-weighted stellar populations are found. Analysis of the WHAN and WHaD diagrams reveal that in fact the central region is retired while the rest of the disc is star forming. We conclude that W2246f is a nice example of a cLIER galaxy, where the central kpc is dominated by old, metal-poor stars with little star formation. The central LIER emission is primarily powered by hot evolved stars, while the rest of the disc displays ongoing star formation. These findings are consistent with a scenario of inside-out quenching.

astro-ph.GA

Hidden Monsters with SPHEREx I: A goldmine for heavily reddened quasars at cosmic noon

Heavily reddened quasars (HRQs) are luminous, dust-obscured broad-line quasars thought to represent a short-lived phase of intense black hole growth and feedback. Previous studies have been limited by small sample sizes, restricting robust statistical analysis. We expand the sample of the most luminous HRQs to enable population-level studies, connecting their spectral energy distributions (SEDs) to other quasar populations and placing them within an evolutionary sequence of massive galaxy and black hole formation. We assemble multiwavelength broadband photometry for the brightest HRQ candidates (K$_{AB}$ < 18 mag) and select AGN with red near-infrared colours (J-K)$_{AB}$ > 1.6. Using SPHEREx spectrophotometry, we confirm HRQs and determine redshifts. Detailed SED fitting allows comparison with other luminous quasars, including a control sample of hyper-luminous, unobscured Quaia quasars and luminous Hot Dust-Obscured Galaxies (Hot DOGs). We confirm 77 new HRQs with redshifts 1.5 < z < 3.9, dust-corrected optical continuum luminosities log$_{10}(\lambda L_\lambda (3000A)$ [erg/s])>47.0, and line-of-sight extinctions 0.4 < E(B-V) < 1.6 (A$_V$ mag). This more than doubles the known HRQs at z > 1.5, including the first seven at z > 3. A UV excess consistent with scattered quasar emission is detected in 76% of HRQs. We show that HRQs are hot-dust poor compared to blue quasars of similar luminosity and redshift. Their 6um continuum luminosities are systematically fainter at fixed 3000A continuum luminosity relative to blue Quaia quasars, indicating deficiency in both hot and warm dust. These results support a scenario in which HRQs represent a blow-out phase, where strong feedback begins clearing obscuring material from central regions.

astro-ph.GA

Heterogeneous Judge-Aware Ranking with Sensitivity, Disagreement, and Confidence

Pairwise comparisons from multiple judges are central to large language model evaluation and preference modeling, yet standard ranking pipelines often pool judgments into a single score vector, treating systematic judge disagreement as noise. We propose Heterogeneous Judge-Aware (HJA) ranking, a structured multi-judge ranking framework that separates consensus ranking, judge-specific sensitivity to consensus, and residual preference disagreement. HJA thereby treats ranking, judge sensitivity, and structured disagreement as separate inferential targets. We establish conditions under which this decomposition is identifiable and develop an anchored alternating algorithm that preserves the identifying geometry. For confidence quantification, we study a fixed-panel repeated-comparison regime in which the judge panel may remain fixed or modest while information grows through repeated judgments. This yields uncertainty statements for consensus and judge-specific ranking contrasts, sensitivity parameters, pairwise probabilities, and summaries of residual disagreement.Experiments on synthetic and real multi-judge comparison data show that HJA improves recovery, robustness, uncertainty calibration, and near-tie performance relative to pooled and sensitivity-only baselines. The fitted model also provides diagnostics for judge disagreement and model-affinity patterns, giving a statistically grounded framework for ranking under heterogeneous comparative judgments.

stat.ME

ParaRNN: An Interpretable and Parallelizable Recurrent Neural Network for Time-Dependent Data

The proliferation of large-scale and structurally complex data has spurred the integration of machine learning methods into statistical modeling. Recurrent neural networks (RNNs), a foundational class of models for time-dependent data, can be viewed as nonlinear extensions of classical autoregressive moving average models. Despite their flexibility and empirical success in machine learning, RNNs often suffer from limited interpretability and slow training, which hinders their use in statistics. This paper proposes the Parallelized RNN (ParaRNN), a novel model composed of multiple small recurrent units. ParaRNN admits an additive representation that decouples recurrent dynamics into interpretable components, whose behavior can be characterized through recurrence features. This interpretability enables its applications in nonparametric regression for time-dependent data, while the design also allows efficient parallelization. The approximation capacity and non-asymptotic prediction error bounds in a nonparametric regression setting are established for ParaRNN. Empirical results on three sequential modeling tasks further demonstrate that ParaRNN achieves performance comparable to vanilla RNNs while offering improved interpretability and efficiency.

stat.ML

AGN Variability with Rubin Observatory in the 2030s

AGN variability offers a direct probe of accretion physics, disk structure, and black hole growth, but progress has been limited by sample size, cadence heterogeneity, and photometric systematics. The Vera C. Rubin Observatory Legacy Survey of Space and Time (LSST) will deliver multi-band light curves for millions of AGN, enabling variability studies at a true population scale. We synthesize recent results from the Zwicky Transient Facility (ZTF), which demonstrate that optical variability amplitudes and timescales are primarily regulated by accretion state, with secondary dependence on black hole mass and redshift, and establish the feasibility of survey-driven continuum reverberation mapping. ZTF measurements reveal optical continuum-emitting region sizes that often exceed standard thin disk predictions, implicating diffuse continuum emission from the broad line region as a significant contributor to observed inter-band lags. We evaluate the implications of LSST cadence and survey strategy, particularly the deep drilling fields, for continuum and emission line reverberation mapping, changing-look AGN, extreme variability quasars, and periodic variability searches. Key limitations of broadband photometric variability are identified, including variable emission line contamination, diffuse BLR continuum emission, and cadence-dependent lag recoverability. We argue that realizing LSST's full scientific potential requires community-scale, standardized variability metric pipelines, probabilistic classification integrated with alert brokers for follow-up triggering, and complementary medium-band photometric observations to isolate the accretion disk continuum. Together, these elements will enable LSST to convert photometric variability into quantitative constraints on accretion disks, BLR structure, and supermassive black hole growth across cosmic time.

astro-ph.GA

Evidence of Gas Depletion in Quasars with Moderate Radio Emission

The energy released by active galactic nuclei (AGNs) is considered to have a profound impact on the cold gas properties of their host galaxies, potentially heating or removing the gas and further suppressing star formation. To understand the feedback from AGN radio activity, we investigate its impacts on the cold gas reservoirs in AGNs with different radio activity levels. We construct a quasar sample with a mean $z\sim1.5$ and a mean $L_{\rm bol}\sim10^{45.8}\ \rm erg\ s^{-1}$, all with Herschel detections to enable estimates of the total gas mass through the galactic dust continuum emission. The sample is then cross-matched with radio catalogs and divided into radio loud (RL) quasars, radio-detected radio quiet (RQ) quasars and radio-undetected quasars based on their radio loudness. Through spectral energy distribution (SED) fitting, we find the radio-detected RQ quasars exhibit evidence of gas deficiency with host galaxies possessing $\sim 0.3$ dex lower dust and gas masses compared to the other two groups, despite being matched in $M_{\rm BH}$, $L_{\rm bol}$, $M_{*}$ and SFR. Furthermore, evidence from optical spectra shows that both the fraction and velocity of outflows are higher in the radio-detected RQ group, suggesting a connection between the ionized gas outflows and the moderate radio activity. These results suggest that the AGN feedback could be more efficient in AGNs with weak/moderate radio emission than in those without radio detection or those with strong radio emission. Further high-resolution observations are needed to understand the interaction between the interstellar medium and the weak/moderate AGN radio activity.

astro-ph.GA

Reduced-Rank Network Autoregression with Grouped Edge Effects

We propose the Edge-Grouped Reduced-Rank Network Autoregressive model (Edge-Grouped RRNAR) for multivariate time series observed over a network. The model allows transmission effects to vary across prespecified and economically interpretable groups of edges, while using a low-rank structure to capture cross-variable dynamics. This structure separates where network transmission occurs from which variables transmit and respond. We develop a topology-aware and block-specific scaled gradient descent algorithm with a convex low-rank initialization, and establish its local linear convergence and non-asymptotic estimation rates. We further derive asymptotic normality for the estimated transition matrix and normalized grouped network coefficients, enabling Wald tests for network relevance and edge-group homogeneity. Simulations demonstrate the finite-sample performance of the proposed estimation and inference procedures. An application to quarterly U.S. industry data linked by the production network reveals distinct predictive transmission intensities across economically defined edge groups, identifies interpretable real-activity and price-cost channels, and improves forecasting relative to standard network and matrix autoregressive models.

stat.ME

Making Wide Stripes Practical: Cascaded Parity LRCs for Efficient Repair and High Reliability

Erasure coding with wide stripes is increasingly adopted to reduce storage overhead in large-scale storage systems. However, existing Locally Repairable Codes (LRCs) exhibit structural limitations in this setting: inflated local groups increase single-node repair cost, multi-node failures frequently trigger expensive global repair, and reliability degrades sharply. We identify a key root cause: local and global parity blocks are designed independently, preventing them from cooperating during repair. We present Cascaded Parity LRCs (CP-LRCs), a new family of wide stripe LRCs that embed structured dependency between parity blocks by decomposing a global parity block across all local parity blocks. This creates a cascaded parity group that preserves MDS-level fault tolerance while enabling low-bandwidth single-node and multi-node repairs. We provide a general coefficient-generation framework, develop repair algorithms exploiting cascading, and instantiate the design with CP-Azure and CP-Uniform. Evaluations on Alibaba Cloud show reductions in repair time of up to 41% for single-node failures and 26% for two-node failures.

cs.DC

Predicting Quasar Counts Detectable in the LSST Survey

The Legacy Survey of Space and Time (LSST), being conducted by the Vera C. Rubin Observatory, is a wide-field multi-band survey that will revolutionize our understanding of extragalactic sources through its unprecedented combination of area and depth. While the LSST survey strategy is still being finalized, the Rubin Observatory team has generated a series of survey simulations using the LSST Operations Simulator to explore the optimal survey strategy that best accommodates the majority of scientific goals. In this study, we utilize the latest simulated data to predict the number of detectable quasars by LSST in each band and evaluate the impact of different survey strategies. We find that the number of quasars and lower luminosity AGNs detected in the baseline strategy (v4.3.1) in the redshift range z=0.3-6.7 will be highest in the i-band and lowest in the u-band. Over 70% of quasars are expected to be detected within the first year in all bands, as LSST will have already reached the break of the luminosity function at most redshifts. With a limiting magnitude of 25.7 mag, we expect to detect 184 million AGNs in the z-band over the 10-year survey, with quasars constituting only 6% of the total AGNs in each band. This arises because, considering that the luminosities of most low-luminosity AGNs are affected by contamination from their host galaxies, we set a magnitude threshold when predicting the number of quasars. We find that variations in the u-band strategy can impact the number of quasar detections. Specifically, the difference between the baseline strategy and that with the largest total exposure in u is 15%. In contrast, changes in rolling strategies, DDF strategies, weather conditions, and Target of Opportunity observations result in variations below 2%. These results provide valuable insights for optimizing approaches to maximize the scientific output of quasar studies.

astro-ph.GA

High-dimensional Autoregressive Modeling for Time Series with Hierarchical Structures

Modern applications have made ubiquitous high-dimensional data, especially time-dependent data, with more and more complicated structures, and it also has become more frequent to encounter the scenario of hierarchical relationships among variables. However, there is still a lack of supervised learning tool in the literature for them. To fill this gap, we introduce a new model-designing framework, and it then combines with unsupervised factor modeling tools to form an efficient and interpretable autoregressive model for high-dimensional time series with hierarchical structures. An ordinary least squares estimation is considered, and its non-asymptotic properties are established. Moreover, we propose an algorithm to search for estimates, and a boosting method is also suggested for hyperparameter selection. Simulation experiments are conducted to evaluate finite-sample performance of the proposed methodology, and its usefulness is demonstrated by an application to the Personality-120 dataset.

stat.ME

Investigating the Impacts of AGN Activities on Dwarf Galaxies with FAST HI Observations

We present the results of Hi line observations towards 26 Active Galactic Nuclei (AGN)-hosting and one star-forming dwarf galaxies (Mstar < 10^9.5 Msun) with the 19-beam spectral line receiver of FAST at 1.4 GHz. Our FAST observed targets are combined with other AGN-hosting dwarf galaxies covered in the ALFALFA footprint to form a more comprehensive sample. Utilizing the information from optical surveys, we further divide them into isolated and accompanied subsamples by their vicinity of nearby massive galaxies. We compare the Hi gas abundance and star-forming rate (SFR) between the subsamples to assess the role of internal and external processes that may regulate the gas content in dwarf galaxies. As a result, we find that AGN are more commonly identified in accompanied dwarf galaxies than in their isolated counterparts. Meanwhile, AGN-hosting dwarf galaxies have slightly but significant lower Hi mass fraction relatively to the non-AGN control sample in accompanied dwarf galaxies. On the other hand, we find a decreasing SFR in AGN-hosting dwarf galaxies towards denser environments, as well as an extremely low incidence of quenched isolated dwarfs within both AGN and non-AGN subsamples. These results indicate that although these AGN could potentially regulate the gas reservoir of dwarf galaxies, environmental effects are likely the dominant quenching mechanism in the low-mass universe.

astro-ph.GA

Improving time series estimation and prediction via transfer learning

There are many time series in the literature with high dimension yet limited sample sizes, such as macroeconomic variables, and it is almost impossible to obtain efficient estimation and accurate prediction by using the corresponding datasets themselves. This paper fills the gap by introducing a novel representation-based transfer learning framework for vector autoregressive models, and information from related source datasets with rich observations can be leveraged to enhance estimation efficiency through representation learning. A two-stage regularized estimation procedure is proposed with well established non-asymptotic properties, and algorithms with alternating updates are suggested to search for the estimates. Our transfer learning framework can handle time series with varying sample sizes and asynchronous starting and/or ending time points, thereby offering remarkable flexibility in integrating information from diverse datasets. Simulation experiments are conducted to evaluate the finite-sample performance of the proposed methodology, and its usefulness is demonstrated by an empirical analysis on 20 macroeconomic variables from Japan and another nine countries.

stat.ME

High-dimensional low-rank matrix regression with unknown latent structures

We study low-rank matrix regression in settings where matrix-valued predictors and scalar responses are observed across multiple individuals. Rather than assuming a fully homogeneous coefficient matrices across individuals, we accommodate shared low-dimensional structure alongside individual-specific deviations. To this end, we introduce a tensor-structured homogeneity pursuit framework, wherein each coefficient matrix is represented as a product of shared low-rank subspaces and individualized low-rank loadings. We propose a scalable estimation procedure based on scaled gradient descent, and establish non-asymptotic bounds demonstrating that the proposed estimator attains improved convergence rates by leveraging shared information while preserving individual-specific signals. The framework is further extended to incorporate scaled hard thresholding for recovering sparse latent structures, with theoretical guarantees in both linear and generalized linear model settings. Our approach provides a principled middle ground between fully pooled and fully separate analyses, achieving strong theoretical performance, computational tractability, and interpretability in high-dimensional multi-individual matrix regression problems.

stat.ME

Probing the Physics of Dusty Outflows through Complex Organic Molecules in the Early Universe

Galaxy-scale outflows are of critical importance for galaxy formation and evolution. Dust grains are the main sites for the formation of molecules needed for star formation but are also important for the acceleration of outflows that can remove the gas reservoir critical for stellar mass growth. Using the MIRI medium-resolution integral field spectrograph aboard the James Webb Space Telescope (JWST), we detect the 3.28 $\mu$m aromatic and the 3.4 $\mu$m aliphatic hydrocarbon dust features in absorption in a redshift 4.601 hot dust-obscured galaxy, blue-shifted by $\Delta$V=$-5250^{+276}_{-339}$ kms$^{-1}$ from the systemic redshift of the galaxy. The extremely high velocity of the dust indicates that the wind was accelerated by radiation pressure from the central quasar. These results pave a novel way for probing the physics of dusty outflows in active galaxies at early cosmic time.

astro-ph.GA

Optimal Repair of $(k+2, k, 2)$ MDS Array Codes

Maximum distance separable (MDS) codes are widely used in distributed storage systems as they provide optimal fault tolerance for a given amount of storage overhead. The seminal work of Dimakis~\emph{et al.} first established a lower bound on the repair bandwidth for a single failed node of MDS codes, known as the \emph{cut-set bound}. MDS codes that achieve this bound are called minimum storage regenerating (MSR) codes. Numerous constructions and theoretical analyses of MSR codes reveal that they typically require exponentially large sub-packetization levels, leading to significant disk I/O overhead. To mitigate this issue, many studies explore the trade-offs between the sub-packetization level and repair bandwidth, achieving reduced sub-packetization at the cost of suboptimal repair bandwidth. Despite these advances, the fundamental question of determining the minimum repair bandwidth for a single failure of MDS codes with fixed sub-packetization remains open. In this paper, we address this challenge for the case of two parity nodes ($n-k=2$) and sub-packetization $\ell=2$. Under these parameters, we establish a correspondence between repair schemes and point sets on the projective line $\mathbb{P}^1$, and then derive a lower bound on repair bandwidth utilizing the sharply 3-transitive action of $\text{PGL}_2(\Fq)$. Furthermore, we extend this lower bound to the repair I/O, and construct two classes of explicit MDS array codes that achieve these bounds, offering practical code designs with provable repair efficiency.

cs.IT