arXiv ScienceSearch

EXPLORE CONNECTIONS

Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation

Follow the relationships that help you find your next source.

Connections are not prepared for this record yet, or its source metadata has changed. The original record remains available while background snapshots are rebuilt.