arXiv ScienceSearch

arXiv subjects

Juergen Gall

Publications and source records attributed to Juergen Gall.

2 recordsLinked to original sources

TQD-Track: Temporal Query Denoising for 3D Multi-Object Tracking

Query denoising has become a standard training strategy for DETR-based detectors. Denoising queries, initialized by perturbing ground truths, share similarities with track queries in a typical DETR-based Multi-Object Tracking (MOT) method, warranting exploration of their potential synergy. However, query denoising in existing MOT methods is performed only within a single frame, preventing trackers from learning inter-frame temporal association from the denoising process. To address this issue, we propose TQD-Track, a Temporal Query Denoising (TQD) method tailored for MOT. In our method, denoising queries are initialized from ground truths in the previous frame and then propagated into the current frame in the same way as track queries, serving as additional independent data association candidates. These denoising queries carry temporal information and instance-specific feature representations, effectively emulating and augmenting track queries. Moreover, to simulate various real-world MOT challenges for robust tracking, we introduce several corresponding noise types to generate diverse denoising queries. We analyze the impact of our temporal query denoising for two tracking paradigms, tracking-by-attention and alternating detection and association, demonstrating its generalization. Extensive experiments on the nuScenes and Argoverse~2 datasets demonstrate that our approach consistently enhances different MOT baselines, requiring only modifications in the training process. Code and models are available at https://github.com/yutongy98/TQD-Track.

cs.CV

Post-Training VLMs for Video Mistake Detection

Human mistakes are inevitable when following instructions, yet they can lead to severe consequences. As such, there has been an increased interest in developing methods for detecting mistakes in videos, with current methods mostly focusing on closed-set protocols. While successful in controlled settings, the closed-set assumption limits their wider applicability, as any changes to the task require collecting new data and re-training models. Instead, we argue that mistake detection methods should learn the general concept of a mistake, rather than overfitting to step-specific details. To reflect this, we introduce the Mistake Detection Video Question Answering (MD-VQA) protocol and accompanying benchmark. MD-VQA tests whether methods can discern if a step was executed correctly with respect to its description, for both seen and unseen actions. To address this important challenge, we propose the first video-language-model post-training technique for mistake detection. Our method uses a tailored reward function to encourage the model to identify discrepancies between an instruction and the corresponding video. Extensive evaluations demonstrate that this approach outperforms zero-shot, supervised fine-tuning, and post-training baselines. Notably, our method generalizes especially well to unseen procedures, for instance, with an improvement of up to 11.6% over the best-performing baseline on EP-VQA, paving the way toward general mistake detection. We release our code and benchmark at https://github.com/FedeSpu/mstk.

cs.CV