arXiv ScienceSearch

arXiv subjects

Jing Zhou

Publications and source records attributed to Jing Zhou.

At least 19 recordsLinked to original sources

Continuous Token-Level Spatio-Temporal Context Modeling for Visual Object Tracking

Spatio-temporal context has become increasingly crucial for visual tracking. However, most existing approaches extract spatio-temporal cues via discrete sampling strategies, which inherently deviate from the continuity of spatio-temporal context, thereby deteriorating tracking performance. To address this challenge, we propose TLCTrack, a novel tracking framework that models token-level spatio-temporal context through continuously updated salient tokens, enabling more accurate target representation. Specifically, TLCTrack incorporates three components: Masked Unidirectional Attention (MUA), Spatial Salient Token Collection (SSTC), and Temporal Salient Token Bank (TSTB) modules. By explicitly integrating spatio-temporal context, MUA extracts discriminative targetaware spatial features in the search region. To avoid the negative impact of background on feature learning, SSTC progressively suppresses background interference, thereby enhancing target spatial representation. Finally, TSTB captures high-quality spatio-temporal information through continuous salient token updates. Extensive experiments on five benchmarks demonstrate that our method achieves superior performance over state-of-the-art trackers. Code and models are available at https://github.com/xiading123/TLCTrack.

cs.CV

A Unified Approach to Interpretable Causal Root Cause Attribution

Understanding why a target metric changes is a fundamental problem in data-driven decision making, beyond anomaly detection alone. We study root cause attribution for metric changes in complex e-commerce systems, focusing on trade-offs between interpretability, efficiency, and causal validity. As a starting point, we extend a metric-decomposition method into a recursive metric-tree framework for multi-level root cause analysis, but this relies on independence and decomposability assumptions that miss complex causal dependencies. In contrast, graphical causal models (GCMs) relax these assumptions and improve causal validity, at the cost of interpretability, higher computational and data demands, and potential attribution target misalignment. Through real-world applications, mathematical proofs, and simulations, we characterize the fundamental sources of these trade-offs. Guided by these insights, we propose a unified, causally informed attribution approach that integrates structural causal information into the metric-tree decomposition framework and corrects key sources of misalignment in GCM-based causal attributions, substantially improving causal validity while preserving interpretability and fast computation. Analytical proofs and simulations demonstrate that the proposed approach produces more accurate root cause attributions, and we also present a real-world application.

stat.ME

A Self-Triggered Agentic Push Recommendation System

Push notification is a critical recommendation scenario on large-scale platforms, allowing the system to proactively reach users outside the application to improve long-term re-engagement. However, designing an optimal push system requires handling a complex action space for the "whether and when" delivery problem under strict system resource constraints. Existing solutions typically fall into two passive paradigms: pre-planned frequency methods that allocate delivery times via offline modeling, limiting real-time adaptability; and fixed-interval triggering methods that periodically poll the system, creating a strict dilemma between excessive computational overhead and diminished optimal timing capture. Furthermore, such multi-stage frameworks severely suffer from local optima. To overcome these limitations, in this paper, we propose STEPS, a proactive, Self-Triggered End-to-end Agentic Push Recommendation System, which is already fully deployed at Douyin with over 1 billion users. STEPS reformulates push recommendation as a self-triggered agentic process in which the system decides not only whether to send a push, but also when to invoke itself again, thereby forming a closed loop that balances real-time effectiveness and efficiency. Specifically, STEPS consists of two decision transformer-based agents: a planning agent that schedules the next system invocation using a gated ordinal regression method, and an execution agent that decides whether to send a push based on trajectory rewards. Furthermore, we introduce a lightweight filtering agent to both control computational overhead and act as a crucial safeguard against unreasonable planning behaviors. Online A/B testing demonstrates that STEPS significantly increases user active days by 0.2843% and reduces the push permission disablement rate by 1.9089%, while the filtering agent reduces computational overhead by 79.42%.

cs.IR

Free-Running Waveguide-Integrated Single-Photon Avalanche Detectors for Visible Light

Waveguide-integrated single-photon avalanche detectors (SPADs) are essential components of integrated photonics platforms for scalable extreme-low-light applications without the use of cryogenics. Here, we demonstrate an integrated SPAD for visible light operating at room temperature in a free-running mode without gating. The device is based on a doped silicon diode end-fire-coupled to a silicon nitride (SiN) photonic integrated circuit (PIC). We investigate a range of lateral and vertical doping profile designs, and operate the devices with a simple current-mode passive quenching circuit. The optimal device is a laterally-doped p-i-n+ SPAD with a maximum photon detection efficiency (PDE) of 1.95 +/- 0.32% for input light at 685 nm wavelength, when reverse-biased at an excess of 1.5 V beyond the breakdown voltage of 15.0 V. We identify promising avenues for improving device performance, which would enable such integrated SPADs to be an attractive choice for cutting-edge integrated photonics solutions in quantum technologies, low-light imaging, and high-speed communications at visible wavelengths.

physics.optics

DBHN-Net: Dual-Branch Hybrid Neural Network For Low-Complexity Monaural Speech Enhancement

Although artificial neural network (ANN) based speech enhancement (SE) methods demonstrate excellent performance, the high computational complexity and high energy consumption hinder their deployment in practical front-end processing tasks.} Currently, the spiking neural networks (SNNs) have shown potential in reducing power consumption. However, the discrete binary activation and complex spatio-temporal dynamics of SNNs often result in information loss. The current challenge therefore focuses on how to maintain performance and reduce computational complexity. To address this issue, this work propose a Dual-Branch Hybrid Neural (DBHN) Network. 1) In terms of network architecture: A dual-branch network integrating ANN and SNN was designed, where the SNN branch reduces power consumption while the ANN branch addresses information loss; The BandSplit and Time-Frequency (TF) -Mamba modules were developed to simultaneously compress energy consumption and enhance model performance; Spiking Feature Extraction Group (SFEG) and Information Transformation Block (ITB) components were implemented with residual connections to mitigate information loss while further refining feature representations. 2) To facilitate inter-branch information fusion: An Interaction module was designed to promote information exchange at various stages of the dual-branch network; A TF-Cross Attention-Fusion module was designed to perform time-frequency domain fusion of dual-branch information while data-adaptively guiding the SNN branch to retain more critical information. Results show that the proposed model maintains superior performance across three public datasets while achieving an average 7.5 fold reduction in computational complexity compared to baseline models.

cs.SD

TabCF: Distributional Control Function Estimation with Tabular Foundation Models

Instrumental variable (IV) and control function (CF) methods are powerful tools for causal effect estimation in the presence of unmeasured confounding, yet most existing approaches target only mean effects and/or demand substantial fitting and tuning effort. In this paper, we introduce a simple method, TabCF, for control function regression using tabular foundation models, which enables accurate, fast, identification-transparent, and tuning-light causal estimation of distributional quantities, such as interventional means and quantiles; we also propose a copula-based approximation for multivariate outcomes. TabCF performs favorably against representative methods across a broad range of small- to medium-sized synthetic and real data scenarios. The central message is two-fold: for practitioners, it highlights that TabCF is an effective tool for distributional causal inference; for researchers, it suggests that the proposed approach could be considered a strong baseline for future method development. Code is available at https://github.com/GepingChen/TabCF.

stat.ML

Detecting Breast Carcinoma Metastasis on Whole-Slide Images by Partially Subsampled Multiple Instance Learning

Breast cancer is the most prevalent cancer in women worldwide. Histopathology image analysis serves as the gold standard for cancer diagnosis. In this regard, whole-slide imaging (WSI), a revolutionary technology in digital pathology, allows for ultrahigh-resolution tissue analysis. Despite its promise, WSI analysis faces significant computational challenges due to its massive data size and tissue heterogeneity. To address this issue, we present a Gaussian mixture based multiple instance learning (MIL) framework for WSI analysis with partially subsampled instances. Our approach models a WSI as a bag of instances (i.e., randomly cropped sub-images), leveraging a bag-based maximum likelihood estimator (BMLE) to predict metastases. Furthermore, we introduce a subsampling-based maximum likelihood estimator (SMLE) to refine predictions by selectively labeling a subset of instances. Extensive evaluations of the breast carcinoma metastasis prediction demonstrate that BMLE surpasses state-of-the-art methods, while the SMLE further improves the prediction accuracy at both bag and instance levels. We find that our method is fairly robust against various plausible model mis-specifications. Theoretical analyses and simulation studies validate the performance and robustness of our methods.

stat.ME

Hypothesis Testing for Penalized Estimating Equations with Cross-Fitted Covariance Calibration

We study hypothesis testing for penalized estimators in settings where the full marginal distribution of a multivariate response is difficult to specify, such as longitudinal data with correlated measurements or high-dimensional heteroscedastic regression. Assuming that the conditional mean model is correctly specified, we establish that the penalized estimating equations admit a $\sqrt{n}$-consistent solution, even when the working covariance structure is misspecified. Our inferential target is a low-dimensional subvector of parameters associated with the mean model. We show that the resulting test statistic converges to a $\chi^2$ distribution, and that its asymptotic power depends on the nuisance covariance function. To mitigate this dependence, we propose estimating the covariance function via cross-fitting, which provides a calibrated and robust procedure for inference.

stat.ME

Universal scaling laws for dynamical-thermal hysteresis

Dynamic hysteresis, the rate-dependent lagged response of materials to external fields, underpins applications from energy-efficient transformers to gas storage systems. A fundamental yet unresolved question is how the hysteresis loop area $A$ scales with the field sweep rate $R$. Here, we reveal that a competition between the field sweep and thermal fluctuations governs a universal crossover between two scaling regimes: $A - A_0 \propto R^{1/3}$ for $R < R^*$ and $A - A_0 \propto R^{2/3}$ for $R > R^*$, where $A_0$ is the quasi-static area and the crossover rate $R^* \propto T/T_c$ depends on the temperature $T$ and the material's critical temperature $T_c$. We demonstrate these scaling laws universally across experiments of magnetic materials, simulations of Ising and metal-organic framework models, and analytical solutions of a stochastic Langevin equation. This framework not only resolves the long-standing non-universality of reported scaling exponents but also provides a direct design principle for the application of dynamic hysteresis.

cond-mat.stat-mech

HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning

Vision-language models (VLMs) show strong multimodal capabilities but still struggle with fine-grained vision-language reasoning. We find that long chain-of-thought (CoT) reasoning exposes diverse failure modes, including perception, reasoning, knowledge, and hallucination errors, which can compound across intermediate steps. However, most existing vision-language data used for reinforcement learning with verifiable rewards (RLVR) does not involve complex reasoning chains that rely on visual evidence throughout, leaving these weaknesses largely unexposed. We therefore propose HopChain, a scalable framework for synthesizing multi-hop vision-language reasoning data for RLVR training of VLMs. Each synthesized multi-hop query forms a logically dependent chain of instance-grounded hops, where earlier hops establish the instances, sets, or conditions needed for later hops, while the final answer remains a specific, unambiguous number suitable for verifiable rewards. We train Qwen3.5-35B-A3B and Qwen3.5-397B-A17B under two RLVR settings: the original data alone, and the original data plus HopChain's multi-hop data, and compare them across 24 benchmarks spanning STEM and Puzzle, General VQA, Text Recognition and Document Understanding, and Video Understanding. Although this multi-hop data is not synthesized for any specific benchmark, it improves 20 of 24 benchmarks on both models, indicating broad and generalizable gains. Consistently, replacing full chained queries with half-multi-hop or single-hop variants reduces the average score across five representative benchmarks from 70.4 to 66.7 and 64.3, respectively. Notably, multi-hop gains peak in long-CoT vision-language reasoning, exceeding 50 points in the ultra-long-CoT regime. These experiments establish HopChain as an effective, scalable framework for synthesizing multi-hop data that improves generalizable vision-language reasoning.

cs.CV

GLEAM: A Multimodal Imaging Dataset and HAMM for Glaucoma Classification

We propose glaucoma lesion evaluation and analysis with multimodal imaging (GLEAM), the first publicly available tri-modal glaucoma dataset comprising scanning laser ophthalmoscopy fundus images, circumpapillary OCT images, and visual field pattern deviation maps, annotated with four disease stages, enabling effective exploitation of multimodal complementary information and facilitating accurate diagnosis and treatment across disease stages. To effectively integrate cross-modal information, we propose hierarchical attentive masked modeling (HAMM) for multimodal glaucoma classification. Our framework employs hierarchical attentive encoders and light decoders to focus cross-modal representation learning on the encoder.

eess.IV

Revealing Behavioral Plasticity in Large Language Models: A Token-Conditional Perspective

In this work, we reveal that Large Language Models (LLMs) possess intrinsic behavioral plasticity-akin to chameleons adapting their coloration to environmental cues-that can be exposed through token-conditional generation and stabilized via reinforcement learning. Specifically, by conditioning generation on carefully selected token prefixes sampled from responses exhibiting desired behaviors, LLMs seamlessly adapt their behavioral modes at inference time (e.g., switching from step-by-step reasoning to direct answering) without retraining. Based on this insight, we propose Token-Conditioned Reinforcement Learning (ToCoRL), a principled framework that leverages RL to internalize this chameleon-like plasticity, transforming transient inference-time adaptations into stable and learnable behavioral patterns. ToCoRL guides exploration with token-conditional generation and keep enhancing exploitation, enabling emergence of appropriate behaviors. Extensive experiments show that ToCoRL enables precise behavioral control without capability degradation. Notably, we show that large reasoning models, while performing strongly on complex mathematics, can be effectively adapted to excel at factual question answering, which was a capability previously hindered by their step-by-step reasoning patterns.

cs.CL

New reformulations for 0-1 quadratic programming problem using quadratic nonconvex reformulation techniques and valid inequalities

It is well-known that the quadratic convex reformulation (QCR) technique can speed up some general-purpose solvers such as CPLEX and Gurobi. Recently, the method of quadratic nonconvex reformulation (QNR) was proposed, which provides an alternative way for accelerating a solver via reformulation technique. This paper proposes several new reformulations for 0-1 quadratic programming problems using the QNR technique. Such a technique provides more flexibility in adding nonconvex quadratic constraints into the problem formulation, so that some valid inequalities, such as the triangle inequalities, can be incorporated into the formulation to tighten the lower bound of the problem. We analyze the effects of the proposed reformulations on the lower bounds implemented in the solver, and propose some methods to maximize the McCormick relaxation bounds of the reformulations. Our numerical experiments compare the proposed reformulations with the existing quadratic convex reformulations, showing the effectiveness of the proposed reformulations on 0-1 quadratic programming problems.

math.OC

A Bayes-Motivated Quadratic-Form Test for High-Dimensional Mean Testing

We propose a two-sample mean test based on the Bayes factor with non-informative priors, specifically designed for scenarios where the dimension $p$ grows with the sample size $n$ with a linear rate $p/n \to c_1 \in (0, \infty)$. We establish the asymptotic normality of the test statistic and the asymptotic power. Through extensive simulations, we demonstrate that the proposed test performs competitively against several existing methods, particularly when the marginal variances of the individual features are heterogeneous and when the sample size is small. Furthermore, our test remains robust under distribution misspecification. The proposed method not only effectively detects both sparse and non-sparse differences in mean vectors but also maintains a well-controlled type I error rate, even in small-sample scenarios. We also demonstrate the performance of our proposed test using the small round blue cell tumors (SRBCT) dataset.

stat.ME

Critical Density-Wave Vestigial Phases of Commensurate Pair Density Wave

The pair-density-wave (PDW) is an exotic pairing state hosting a spatially modulated pairing order parameter, which has attracted great interest. Due to its simultaneously breaking U(1)-gauge and translational symmetries, intriguing vestigial phases which restore only one broken symmetry can emerge at an intermediate temperature regime. Previously, investigations on the vestigial phases of PDW were mainly focused on incommensurate PDW. However, the experimentally observed PDW is usually commensurate, whose vestigial phases have not been systematically investigated. Here we study the vestigial phases of 2D commensurate PDW with $n$-times expanded unit vectors, hosting different numbers of wave vectors. Based on the Ginzburg-Landau theory, we get the low energy effective model Hamiltonian. Subsequent renormalization group (RG) and Monte-Carlo (MC) studies are conducted to obtain the phase diagram and spatial dependent correlation functions. Our RG and MC calculations consistently yield the following result. For $n\le 4$, besides the charge-4e/2e superconductivity, there exists the translational symmetry broken charge-density-wave (CDW) vetigial phase. Intriguingly, for $n\ge 5$, the restore of the translational symmetry with increasing temperature is realized through two successive Berezinskii-Kosterlitz-Thouless transitions. Such a two-step process leads into two critical vestigial phases, i.e. the critical-PDW and the critical-CDW phases, in which the discrete translational symmetry is quasily broken, leading into a power-law decaying density-density correlation even at 2D. Our work appeals for experimental verifications.

cond-mat.str-el

On angular dependent response to gravitational-wave signals for time-delay interferometry combinations

Space-based gravitational wave (GW) detectors are designed for wave sources in the millihertz band with different locations and orientations. Time-delay interferometry (TDI) technique is an indispensable ingredient in space-borne GW detection that effectively suppresses the laser phase noise. The abundant TDI solutions derived in the literature also feature distinct angular-dependent sensitivities. Because a GW source's angular location is unknown prior to the signals' detection, a solid-angle average is often performed when analyzing the sensitivity function of a given TDI combination. The present study explores the angular dependence of the detector's sensitivity. This detail is relevant, because once the initial detection is achieved, the source's location can be extracted and used to provide information on a refined TDI combination tailored for the specific GW source. As the TDI technique is a post-processing algorithm, such a procedure can be implemented in practice. We evaluate the angular dependence of the detector's response function to the GW signals for different TDI combinations as a function of the orientation angles. Moreover, we classify the response functions into seven categories at the low-frequency limit, leveraging the characteristics of the underlying geometrical TDI combinations. By further averaging out the azimuthal angle $\phi_D$ in the detector's plane, the main features of the resulting response functions and their zenithal dependence with respect to the GW source are scrutinized. The findings presented in this work provide pertinent insights for ongoing space-borne detector programs.

gr-qc

Qwen3-VL Technical Report

We introduce Qwen3-VL, the most capable vision-language model in the Qwen series to date, achieving superior performance across a broad range of multimodal benchmarks. It natively supports interleaved contexts of up to 256K tokens, seamlessly integrating text, images, and video. The model family includes both dense (2B/4B/8B/32B) and mixture-of-experts (30B-A3B/235B-A22B) variants to accommodate diverse latency-quality trade-offs. Qwen3-VL delivers three core pillars: (i) markedly stronger pure-text understanding, surpassing comparable text-only backbones in several cases; (ii) robust long-context comprehension with a native 256K-token window for both text and interleaved multimodal inputs, enabling faithful retention, retrieval, and cross-referencing across long documents and videos; and (iii) advanced multimodal reasoning across single-image, multi-image, and video tasks, demonstrating leading performance on comprehensive evaluations such as MMMU and visual-math benchmarks (e.g., MathVista and MathVision). Architecturally, we introduce three key upgrades: (i) an enhanced interleaved-MRoPE for stronger spatial-temporal modeling across images and video; (ii) DeepStack integration, which effectively leverages multi-level ViT features to tighten vision-language alignment; and (iii) text-based time alignment for video, evolving from T-RoPE to explicit textual timestamp alignment for more precise temporal grounding. Under comparable token budgets and latency constraints, Qwen3-VL achieves superior performance in both dense and Mixture-of-Experts (MoE) architectures. We envision Qwen3-VL serving as a foundational engine for image-grounded reasoning, agentic decision-making, and multimodal code intelligence in real-world workflows.

cs.CV

The evolution of CH in Planck Galactic Cold Clumps

Methylidyne (CH) has long been considered a reliable tracer of molecular gas in the low-to-intermediate extinction range. Although extended CH 3.3 GHz emission is commonly observed in diffuse and translucent clouds, observations in cold, dense clumps are rare. In this work, we conducted high-sensitivity CH observations toward 27 PGCCs with the Arecibo 305m telescope. Toward each source, the CH data were analyzed in conjunction with $^{13}$CO (1--0), HINSA, and H$_2$ column densities. Our results revealed ubiquitous subsonic velocity dispersions of CH, in contrast to $^{13}$CO, which is predominantly supersonic. The findings suggest that subsonic CH emissions may trace dense, low-turbulent gas structures in PGCCs. To investigate environmental effects, particularly the cosmic-ray ionization rate (CRIR), we estimated CRIR upper limits from HINSA, yielding values from $(8.1\pm4.7)\times10^{-18}$ to $(2.0\pm0.8)\times10^{-16}$ s$^{-1}$ ($N_{H_2}$ from $(1.7\pm0.2)\times10^{21}$ to $(3.6\pm0.4)\times10^{22}$~cm$^{-2}$). This result favors theoretical predictions of a cosmic-ray attenuation model, in which the interstellar spectra of low-energy CR protons and electrons match {\it Voyager} measurements, although alternative models cannot yet be ruled out. The abundance of CH decreases with increasing column density, while showing a positive dependence on the CRIR, which requires atomic oxygen not heavily depleted to dominate CH destruction in PGCCs. By fitting the abundance of CH with an analytic formula, we place constraints on atomic O abundance ($2.4\pm0.4\times10^{-4}$ with respect to total H) and C$^+$ abundance ($7.4\pm0.7\times10^{13}\zeta_2/n_{\rm H_2}$). These findings indicate that CH formation is closely linked to the C$^+$ abundance, regulated by cosmic-ray ionization, while other processes, such as turbulent diffusive transport, might also contribute a non-negligible effect.

astro-ph.GA