arXiv ScienceSearch

arXiv subjects

Bin Zhong

Publications and source records attributed to Bin Zhong.

11 recordsLinked to original sources

Enhanced evidence of $X(7200)$ and improved measurements of $X(6900)$ parameters from a combined LHCb-ATLAS-CMS analysis

We report stronger evidence for the $X(7200)$ state and markedly improved measurements of the $X(6900)$ resonance parameters based on a combined analysis of the di-$J/\psi$ mass spectrum using published data from LHCb, ATLAS, and CMS. Through simultaneous fits to the datasets from all three experiments, we observe the $X(6900)$ with overwhelming significance ($>12\sigma$) and determine its mass and width with improved precision. For the $X(7200)$, we find consistent signals across multiple interference models, with significances ranging from $3.7\sigma$ to $6.6\sigma$; in the best-fit model (the CMS three-resonance scheme), the significance reaches $6.6\sigma$, providing substantially stronger evidence for this state. Our results underscore the essential role of interference effects in fully charmed tetraquark spectroscopy and offer new constraints on their production mechanisms at the LHC.

hep-ex

OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios

Spatio-Temporal Video Grounding (STVG) aims to localize target objects in videos based on natural language descriptions. Despite recent advances in Multimodal Large Language Models, a significant gap remains between current models and real-world demands involving diverse objects and complex queries. We attribute this to limited benchmark scope, causing models to exhibit category bias, oversimplified reasoning, and poor linguistic robustness. To address these limitations, we introduce OmniGround, a comprehensive benchmark with 3,475 videos spanning 81 categories and complex real-world queries. We propose the Forward-Backward-Refinement annotation pipeline that combines multi-directional tracking with intelligent error correction for high-quality labels. We further introduce DeepSTG, a systematic evaluation framework quantifying dataset quality across four complementary dimensions beyond superficial statistics. Evaluations reveal performance average drop of 10.4% on complex real-world scenes, particularly with small/occluded objects and intricate spatial relations. Motivated by these, we propose PG-TAF, a training-free two-stage framework decomposing STVG into high-level temporal grounding and fine-grained spatio-temporal propagation. Experiments demonstrate PG-TAF achieves 25.6% and 35.6% improvements in m\_tIoU and m\_vIoU on OmniGround with consistent gains across four benchmarks.

cs.CV

Agentic Video Intelligence: A Flexible Framework for Advanced Video Exploration and Understanding

Video understanding requires not only visual recognition but also complex reasoning. While Vision-Language Models (VLMs) demonstrate impressive capabilities, they typically process videos largely in a single-pass manner with limited support for evidence revisit and iterative refinement. While recently emerging agent-based methods enable long-horizon reasoning, they either depend heavily on expensive proprietary models or require extensive agentic RL training. To overcome these limitations, we propose Agentic Video Intelligence (AVI), a flexible and training-free framework that can mirror human video comprehension through system-level design and optimization. AVI introduces three key innovations: (1) a human-inspired three-phase reasoning process (Retrieve-Perceive-Review) that ensures both sufficient global exploration and focused local analysis, (2) a structured video knowledge base organized through entity graphs, along with multi-granularity integrated tools, constituting the agent's interaction environment, and (3) an open-source model ensemble combining reasoning LLMs with lightweight base CV models and VLM, eliminating dependence on proprietary APIs or RL training. Experiments on LVBench, VideoMME-Long, LongVideoBench, and Charades-STA demonstrate that AVI achieves competitive performance while offering superior interpretability.

cs.CV

APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval

Current multimodal large language models (MLLMs) struggle with hour-level video understanding, facing significant challenges not only in modeling the substantial information volume of long videos but also in overcoming the memory wall and resource constraints during both training and inference. Although recent training-free approaches have alleviated resource demands by compressing visual features, their reliance on incomplete visual information limits the performance potential. To address these limitations, we propose Adaptive Pivot Visual information Retrieval (APVR), a training-free framework that hierarchically retrieves and retains sufficient and important visual information. It breakthroughs the memory wall limitation via two complementary components: Pivot Frame Retrieval employs query expansion and iterative spatio-semantic confidence scoring to identify relevant video frames, and Pivot Token Retrieval performs query-aware attention-driven token selection within up to 1024 pivot frames. This dual granularity approach enables the processing of hour-long videos while maintaining semantic fidelity. Experimental validations on three different baseline MLLMs demonstrate significant performance improvements up to 9.5\%, 4.6\% and 9.7\% on LongVideoBench, VideoMME and MLVU, respectively. APVR achieves state-of-the-art results for both training-free and training-based approaches.

cs.CV

Close-range Human Following Control on a Cane-type Robot with Multi-camera Fusion

Cane-type robots have been utilized to assist and supervise the mobility-impaired population. One essential technique for cane-type robots is human following control, which allows the robot to follow the user. However, the limited perceptible information of humans by sensors at close range, combined with the occlusion caused by lower limb swing during normal walking, affect the localization of users. These limitations make it difficult to achieve human following at close range.To address these challenges, this study developed a new cane-type wheeled robot and proposed a novel human-following control with multi-camera fusion. This control system mainly consists of two parts: 1) a human following controller that locates a user by multi-camera fusion and generates control signals to follow the user. 2) a cane robot controller designed to steer the cane robot to a target position. The proposed strategy's effectiveness has been validated in outdoor experiments with six healthy subjects. The experimental scenarios included different terrains (i.e., straight, turning, and inclined paths), road conditions (i.e., flat and rough roads), and walking speeds. The obtained results showed that the average tracking error for position and orientation was less than 5 cm and 15{\deg} respectively across all scenarios. Moreover, the cane robot can effectively adapt to a wide range of individual gait patterns and achieve stable human following at daily walking speeds (0.75 m/s - 1.45 m/s).

eess.SY

CLS: Cross Labeling Supervision for Semi-Supervised Learning

It is well known that the success of deep neural networks is greatly attributed to large-scale labeled datasets. However, it can be extremely time-consuming and laborious to collect sufficient high-quality labeled data in most practical applications. Semi-supervised learning (SSL) provides an effective solution to reduce the cost of labeling by simultaneously leveraging both labeled and unlabeled data. In this work, we present Cross Labeling Supervision (CLS), a framework that generalizes the typical pseudo-labeling process. Based on FixMatch, where a pseudo label is generated from a weakly-augmented sample to teach the prediction on a strong augmentation of the same input sample, CLS allows the creation of both pseudo and complementary labels to support both positive and negative learning. To mitigate the confirmation bias of self-labeling and boost the tolerance to false labels, two different initialized networks with the same structure are trained simultaneously. Each network utilizes high-confidence labels from the other network as additional supervision signals. During the label generation phase, adaptive sample weights are assigned to artificial labels according to their prediction confidence. The sample weight plays two roles: quantify the generated labels' quality and reduce the disruption of inaccurate labels on network training. Experimental results on the semi-supervised classification task show that our framework outperforms existing approaches by large margins on the CIFAR-10 and CIFAR-100 datasets.

cs.CV

An Iterative Weighting Method to Apply ISR Correction to $e^+e^-$ Hadronic Cross-section Measurements

Initial state radiation (ISR) plays an important role in $e^+$$e^-$ collision experiments such as the BESIII. To correct the ISR effects in measurements of hadronic cross-sections of $e^+e^-$ annihilation, an iterative method that weights simulated ISR events is proposed here to assess the efficiency of event selection and the ISR correction factor for the observed cross-section. The simulated ISR events were generated only once, and the obtained cross-sectional line shape was used iteratively to weigh the same simulated ISR events to evaluate the efficiency and corrections until the results converge. Compared with the method of generating ISR events iteratively, the proposed weighting method provides consistent results, and reduces the computational time and disk space required by a factor of five or more, thus speeding-up $e^+e^-$ hadronic cross-section measurements.

hep-ex

Weak Supervision for Fake News Detection via Reinforcement Learning

Today social media has become the primary source for news. Via social media platforms, fake news travel at unprecedented speeds, reach global audiences and put users and communities at great risk. Therefore, it is extremely important to detect fake news as early as possible. Recently, deep learning based approaches have shown improved performance in fake news detection. However, the training of such models requires a large amount of labeled data, but manual annotation is time-consuming and expensive. Moreover, due to the dynamic nature of news, annotated samples may become outdated quickly and cannot represent the news articles on newly emerged events. Therefore, how to obtain fresh and high-quality labeled samples is the major challenge in employing deep learning models for fake news detection. In order to tackle this challenge, we propose a reinforced weakly-supervised fake news detection framework, i.e., WeFEND, which can leverage users' reports as weak supervision to enlarge the amount of training data for fake news detection. The proposed framework consists of three main components: the annotator, the reinforced selector and the fake news detector. The annotator can automatically assign weak labels for unlabeled news based on users' reports. The reinforced selector using reinforcement learning techniques chooses high-quality samples from the weakly labeled data and filters out those low-quality ones that may degrade the detector's prediction performance. The fake news detector aims to identify fake news based on the news content. We tested the proposed framework on a large collection of news articles published via WeChat official accounts and associated user reports. Extensive experiments on this dataset show that the proposed WeFEND model achieves the best performance compared with the state-of-the-art methods.

cs.SI

Transfer Value Iteration Networks

Value iteration networks (VINs) have been demonstrated to have a good generalization ability for reinforcement learning tasks across similar domains. However, based on our experiments, a policy learned by VINs still fail to generalize well on the domain whose action space and feature space are not identical to those in the domain where it is trained. In this paper, we propose a transfer learning approach on top of VINs, termed Transfer VINs (TVINs), such that a learned policy from a source domain can be generalized to a target domain with only limited training data, even if the source domain and the target domain have domain-specific actions and features. We empirically verify that our proposed TVINs outperform VINs when the source and the target domains have similar but not identical action and feature spaces. Furthermore, we show that the performance improvement is consistent across different environments, maze sizes, dataset sizes as well as different values of hyperparameters such as number of iteration and kernel size.

cs.LG

Doubly charmed tetraquarks in a diquark-antidiquark model

We study the spectra of the doubly charmed tetraquark states in a diquark-antidiquark model. The doubly charmed tetraquark states form an antitriplet and a sextet configurations according to flavor SU(3) symmetry. For the tetraquark state $[qq'][\bar c\bar c]$, we show the mass for both bound and excited states. The two-body decays of tetraquark states $T^{cc}[0^+]$ and $T^{cc}[1^{--}]$ to charmed mesons have also been studied. In the end,the doubly charmed tetraquarks decays to a charmed baryon and a light baryon have been studied in the SU(3) flavor symmetry.

hep-ph

$CP$ Invariance Study of $J/\psi\to\Lambda\bar\Lambda$ and $\Lambda$ Nonpleptonic Decays in Helicity Frame

We present the joint helicity amplitudes for $J/\psi \to \Lambda \bar{\Lambda}$, $\Lambda(\bar\Lambda)$ decays to different final states in the helicity frame. Two observables to search for $CP$ violation in $J/\psi\to\Lambda\bar\Lambda$ can be expressed with the information of helicity angles of baryon and antibaryon. Four decay parameters of $\Lambda$ and $\bar\Lambda$, namely, $\alpha_-,\alpha_+,\alpha_0$ and $\bar\alpha_0$, can be obtained with the joint helicity amplitude equations by the likelihood fit method. With the data sample of $10^{10}$ $J/\psi$ decays accumulated by BESIII, the precision of the measurements is estimated to be about $10^{-3}$.

hep-ph