arXiv · 2609.33450
Native Association: Confidence-Aware Human Perception in the Wild with a Foundation VLM
Abstract
Extracting who is where, on which team, wearing which number from a broadcast frame is typically done by stitching a detector, an OCR engine, and classifiers together -- and the stitching step swaps identities under occlusion. We make association native instead: a 0.77B vision-language model (Florence-2) is fine-tuned to emit all per-person attributes as one grammar-constrained sequence, with each attribute generated inside its owner's block. Output is therefore schema-valid on every frame by construction, and no post-hoc binding step exists to attach a correctly read number to the wrong player: residual misassociation is pure perception error, $\approx4\times$ rarer than zero-shot-prompted frontier APIs' (0.057 vs. 0.21-0.24). On a frozen multi-sport test set, this single pass reaches 0.95 detection F1 (APIs: 0.65-0.75). A single extra forward pass yields a per-field confidence that supports a reject option (jersey precision $0.71\rightarrow0.96$ at half coverage) and routes a training-free zoom-and-re-read for small players. Surprisingly, once the grammar is learned, further parameter-efficient tuning yields no measurable gain under the adaptation configurations we test; the identical recipe on WIDER-Attribute reaches 93.1 mAP given-box, yields the first detection-coupled end-to-end results under its standard test protocol (84.5 mAP), and reproduces the same tuning result. In this regime, the gains live in the structure, not in added weights.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Igal Dmitriev, Ofir Liba. 2026-09-27. Native Association: Confidence-Aware Human Perception in the Wild with a Foundation VLM. https://arxiv.org/abs/2609.33450
Cite the original work for its findings. Save a collection to share your selection of sources.