arXiv · 2609.17913
Face-voice Association across LAnguages and Gender (FLAG) 2027 Challenge Evaluation Plan
Abstract
Face--voice association models may rely on language or gender cues in the voice rather than on speaker-specific voice characteristics, which can lead to a performance deterioration when the model has to identify a multilingual speaker or distinguis same-gender speakers. To investigate these issues, we introduce the Face-voice Association across LAnguages and Gender (FLAG) 2027 Challenge. The challenge formulates face--voice association as a cross-modal verification task: given a voice, identify the speaker's face from a ``gallery'' of faces consisting of the speaker's face and a set of negative samples. Models are evaluated on identities not present in the training data (``unseen'') and both for languages present or absent from the training data (``heard'' and ``unheard''). Two evaluation settings are used to test models' reliance on gender: a standard, unconstrained and a gender-constrained one, where the latter uses a same-gender gallery. The performance of existing, baseline models in these settings reveals that models performance degrades under language shifts and in gender-constrained settings, highlighting the need to foster the development of models that capture identity-specific aspects beyond language and gender. The challenge provides a benchmark dataset, pretrained baseline models, and an evaluation framework to advance face--voice association.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Marta Moscati, Swapnil Khandoker, Muhammad Saad Saeed, Shah Nawaz, Fatima Noor, Rohan Kumar Das, Mubashir Noman, Junaid Mir, Muhammad Haroon Yousaf, Khalid Malik, Markus Schedl. 2026-09-15. Face-voice Association across LAnguages and Gender (FLAG) 2027 Challenge Evaluation Plan. https://arxiv.org/abs/2609.17913
Cite the original work for its findings. Save a collection to share your selection of sources.