arXiv ScienceSearch

arXiv subjects

Jianquan Liu

Publications and source records attributed to Jianquan Liu.

13 recordsLinked to original sources

BiMTokenizer: Preserving Semantic-Acoustic Balance in Low-Bitrate Speech Tokenization via Bidirectional State-Space Modeling

Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, recent works increasingly adopt dual-tower architectures to decouple semantic and acoustic modeling with separate encoders. However, these dual-tower designs incur substantial architectural overhead. To avoid such complexity, we revisit the single-tower paradigm and propose BiMTokenizer, a low-bitrate speech codec (around 1.1 kbps) combining a bidirectional state-space backbone with Residual Spherical Leech Quantization (RSLQ). The bidirectional backbone strengthens temporal modeling, while RSLQ offers a fixed, well-separated lattice bottleneck for robust semantic and acoustic tokenization without learned-codebook collapse. Experiments show that BiMTokenizer achieves superior acoustic reconstruction and the lowest WER among low-bitrate codec baselines across both clean and noisy environments, while using less than half the parameters of recent dual-tower baselines. Furthermore, its robust semantic representations yield strong performance on downstream speech understanding tasks, confirming that a well-designed single-tower codec can preserve the semantic-acoustic balance at low bitrates. The code and model weights are available at https://github.com/ZhangXinWhut/BiMTokenizer.

cs.SD

Object-Centric Framework for Video Moment Retrieval

Most existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such representations often fail to capture fine-grained object semantics and appearance, which are crucial for localizing moments described by object-oriented queries involving specific entities and their interactions. In particular, temporal dynamics at the object level have been largely overlooked, limiting the effectiveness of existing approaches in scenarios requiring detailed object-level reasoning. To address this limitation, we propose a novel object-centric framework for moment retrieval. Our method first extracts query-relevant objects using a scene graph parser and then generates scene graphs from video frames to represent these objects and their relationships. Based on the scene graphs, we construct object-level feature sequences that encode rich visual and semantic information. These sequences are processed by a relational tracklet transformer, which models spatio-temporal correlations among objects over time. By explicitly capturing object-level state changes, our framework enables more accurate localization of moments aligned with object-oriented queries. We evaluated our method on three benchmarks: Charades-STA, QVHighlights, and TACoS. Experimental results demonstrate that our method outperforms existing state-of-the-art methods across all benchmarks.

cs.CV

KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding

We propose KFS-Bench, the first benchmark for key frame sampling in long video question answering (QA), featuring multi-scene annotations to enable direct and robust evaluation of sampling strategies. Key frame sampling is crucial for efficient long-form video understanding. In long video QA, selecting informative frames enables multimodal large language models (MLLMs) to improve both accuracy and efficiency. KFS-Bench addresses the limitation of prior works that only indirectly assess frame selection quality via QA accuracy. By providing ground-truth annotations of multiple disjoint scenes required per question, KFS-Bench allows us to directly analyze how different sampling approaches capture essential content across an entire long video. Using KFS-Bench, we conduct a comprehensive study of key frame sampling methods and identify that not only sampling precision but also scene coverage and sampling balance are the key factors influencing QA performance. Regarding all the factors, we design a novel sampling quality metric that correlates with QA accuracy. Furthermore, we develop a novel key frame sampling method that leverages question-video relevance to balance sampling diversity against question-frame similarity, thereby improving coverage of relevant scenes. Our adaptively balanced sampling approach achieves superior performance in both key frame sampling and QA performance. The benchmark is available at https://github.com/NEC-VID/KFS-Bench.

cs.CV

Speaking Clearly: A Simplified Whisper-Based Codec for Low-Bitrate Speech Coding

Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, prevailing methods augment acoustic codecs with complex semantic supervision. We explore the opposite direction: a semantic-first approach that starts from a semantically-capable model and adapts it for high-fidelity acoustic reconstruction. Through empirical analysis, we discover that targeted architectural simplification can unlock the acoustic modeling potential of Whisper, a text-aligned Automatic Speech Recognition (ASR) model. Based on this finding, we propose SimWhisper-Codec, a novel codec that balances the semantic and acoustic preservation by leveraging a frozen, simplified Whisper encoder without requiring external supervision. Experimental results demonstrate that SimWhisper-Codec achieves superior performance in both semantic preservation and acoustic quality compared to semantically-supervised codecs such as Mimi Codec and SpeechTokenizer at similar bitrates, validating the effectiveness of our semantic-first approach. Code is available at https://github.com/ZhangXinWhut/SimWhisper-Codec.

cs.SD

DFQ-ViT: Data-Free Quantization for Vision Transformers without Fine-tuning

Data-Free Quantization (DFQ) enables the quantization of Vision Transformers (ViTs) without requiring access to data, allowing for the deployment of ViTs on devices with limited resources. In DFQ, the quantization model must be calibrated using synthetic samples, making the quality of these synthetic samples crucial. Existing methods fail to fully capture and balance the global and local features within the samples, resulting in limited synthetic data quality. Moreover, we have found that during inference, there is a significant difference in the distributions of intermediate layer activations between the quantized and full-precision models. These issues lead to a severe performance degradation of the quantized model. To address these problems, we propose a pipeline for Data-Free Quantization for Vision Transformers (DFQ-ViT). Specifically, we synthesize samples in order of increasing difficulty, effectively enhancing the quality of synthetic data. During the calibration and inference stage, we introduce the activation correction matrix for the quantized model to align the intermediate layer activations with those of the full-precision model. Extensive experiments demonstrate that DFQ-ViT achieves remarkable superiority over existing DFQ methods and its performance is on par with models quantized through real data. For example, the performance of DeiT-T with 3-bit weights quantization is 4.29% higher than the state-of-the-art. Our method eliminates the need for fine-tuning, which not only reduces computational overhead but also lowers the deployment barriers for edge devices. This characteristic aligns with the principles of Green Learning by improving energy efficiency and facilitating real-world applications in resource-constrained environments.

cs.CV

What is Next when Sequential Prediction Meets Implicitly Hard Interaction?

Hard interaction learning between source sequences and their next targets is challenging, which exists in a myriad of sequential prediction tasks. During the training process, most existing methods focus on explicitly hard interactions caused by wrong responses. However, a model might conduct correct responses by capturing a subset of learnable patterns, which results in implicitly hard interactions with some unlearned patterns. As such, its generalization performance is weakened. The problem gets more serious in sequential prediction due to the interference of substantial similar candidate targets. To this end, we propose a Hardness Aware Interaction Learning framework (HAIL) that mainly consists of two base sequential learning networks and mutual exclusivity distillation (MED). The base networks are initialized differently to learn distinctive view patterns, thus gaining different training experiences. The experiences in the form of the unlikelihood of correct responses are drawn from each other by MED, which provides mutual exclusivity knowledge to figure out implicitly hard interactions. Moreover, we deduce that the unlikelihood essentially introduces additional gradients to push the pattern learning of correct responses. Our framework can be easily extended to more peer base networks. Evaluation is conducted on four datasets covering cyber and physical spaces. The experimental results demonstrate that our framework outperforms several state-of-the-art methods in terms of top-k based metrics.

cs.LG

Basketball Player's Value Evaluation by a Networks-based Variant Parameter Hidden Markov Model

Determining the value of basketball players through analyzing the players' behavior is important for the managers of modern basketball teams. However, conventional methods always utilize isolated statistical data, leading to ineffective and inaccurate evaluations. Existing models based on dynamic network theory offer major improvements to the results of such evaluations, but said models remain imprecise because they focus merely on evaluating the values of individual players rather than considering them within their current teams. To solve this problem, we propose an analysis and evaluation model based on networks and a hidden Markov model. To the best of our knowledge, we are the first to combine a network form representing the players who are playing with the use of a hidden Markov model to mine the network and generate the desired results. Applying our approach to SportVU data collected from the National Basketball Association shows that this analysis and evaluation model can effectively analyze the performance of each player in a game and provides an assistive tool for team managers.

cs.SI

Weakly-Supervised Multi-Person Action Recognition in 360$^{\circ}$ Videos

The recent development of commodity 360$^{\circ}$ cameras have enabled a single video to capture an entire scene, which endows promising potentials in surveillance scenarios. However, research in omnidirectional video analysis has lagged behind the hardware advances. In this work, we address the important problem of action recognition in top-view 360$^{\circ}$ videos. Due to the wide filed-of-view, 360$^{\circ}$ videos usually capture multiple people performing actions at the same time. Furthermore, the appearance of people are deformed. The proposed framework first transforms omnidirectional videos into panoramic videos, then it extracts spatial-temporal features using region-based 3D CNNs for action recognition. We propose a weakly-supervised method based on multi-instance multi-label learning, which trains the model to recognize and localize multiple actions in a video using only video-level action labels as supervision. We perform experiments to quantitatively validate the efficacy of the proposed method and qualitatively demonstrate action localization results. To enable research in this direction, we introduce 360Action, the first omnidirectional video dataset for multi-person action recognition.

cs.CV

Statistical Detection of Collective Data Fraud

Statistical divergence is widely applied in multimedia processing, basically due to regularity and interpretable features displayed in data. However, in a broader range of data realm, these advantages may no longer be feasible, and therefore a more general approach is required. In data detection, statistical divergence can be used as a similarity measurement based on collective features. In this paper, we present a collective detection technique based on statistical divergence. The technique extracts distribution similarities among data collections, and then uses the statistical divergence to detect collective anomalies. Evaluation shows that it is applicable in the real world.

cs.DB

Alternative Awaiting and Broadcast for Two-Way Relay Fading Channels

We investigate a two-way relay (TWR) fading channel based on store-and-forward (SF), where two source nodes wish to exchange information with the help of a relay node. A new upper bound on the ergodic sum-capacity for the TWR fading system is derived when delay tends to infinity.We further propose two alternative awaiting and broadcast (AAB) schemes: pure partial decoding (PPD) with SF-I and combinatorial decoding (CBD) with SF-II, which approach the new upper bound at high SNR with unbounded and bounded delay respectively. Numerical results show that the proposed AAB schemes significantly outperform the traditional physical layer network coding (PLNC) methods without delay. Compared to the traditional TWR schemes without delay, the proposed CBD with SF-II method significantly improves the maximum sum-rate with an average delay of only some dozen seconds in the relay buffer.

cs.IT

Alternative Awaiting and Broadcast for Two-Way Relay Fading Channels

We investigate a two-way relay (TWR) fading channel where two source nodes wish to exchange information with the help of a relay node. Given traditional TWR protocols, transmission rates in both directions are known to be limited by the hop with lower capacity, i.e., the min operations between uplink and downlink. In this paper, we propose a new transmission protocol, named as alternative awaiting and broadcast (AAB), to cancel the min operations in the TWR fading channels. The operational principles, new upper bound on ergodic sum-capacity (ESC) and convergence behavior of average delay of signal transmission (ST) (in relay buffer) for the proposed AAB protocol are analyzed. Moreover, we propose a suboptimal encoding/decoding solution for the AAB protocol and derive an achievable ergodic sum-rate (ESR) with corresponding average delay of ST. Numerical results show that 1) the proposed AAB protocol significantly improves the achievable ESR compared to the traditional TWR protocols, 2) considering the average delay of system service (SS) (in source buffer), the average delay of ST induced by the proposed AAB protocol is very small and negligible.

cs.IT

Pairwise Check Decoding for LDPC Coded Two-Way Relay Block Fading Channels

Partial decoding has the potential to achieve a larger capacity region than full decoding in two-way relay (TWR) channels. Existing partial decoding realizations are however designed for Gaussian channels and with a static physical layer network coding (PLNC). In this paper, we propose a new solution for joint network coding and channel decoding at the relay, called pairwise check decoding (PCD), for low-density parity-check (LDPC) coded TWR system over block fading channels. The main idea is to form a check relationship table (check-relation-tab) for the superimposed LDPC coded packet pair in the multiple access (MA) phase in conjunction with an adaptive PLNC mapping in the broadcast (BC) phase. Using PCD, we then present a partial decoding method, two-stage closest-neighbor clustering with PCD (TS-CNC-PCD), with the aim of minimizing the worst pairwise error probability. Moreover, we propose the minimum correlation optimization (MCO) for selecting the better check-relation-tabs. Simulation results confirm that the proposed TS-CNC-PCD offers a sizable gain over the conventional XOR with belief propagation (BP) in fading channels.

cs.IT

Superimposed XOR: Approaching Capacity Bounds of the Two-Way Relay Channels

In two-way relay channels, bitwise XOR and symbol-level superposition coding are two popular network-coding based relaying schemes. However, neither of them can approach the capacity bound when the channels in the broadcast phase are asymmetric. In this paper, we present a new physical layer network coding (PLNC) scheme, called \emph{superimposed XOR}. The new scheme advances the existing schemes by specifically taking into account the channel asymmetry as well as information asymmetry in the broadcast phase. We obtain its achievable rate regions over Gaussian channels when integrated with two known time control protocols in two-way relaying. We also demonstrate their average maximum sum-rates and service delay performances over fading channels. Numerical results show that the proposed superimposed XOR achieves a larger rate region than both XOR and superposition and performs much better over fading channels. We further deduce the boundary of its achievable rate region of the broadcast phase in an explicit and analytical expression. Based on these results, we then show that the gap to the capacity bound approaches zero at high signal-to-noise ratio.

cs.IT