arXiv ScienceSearch

arXiv subjects

Brian Song

Publications and source records attributed to Brian Song.

2 recordsLinked to original sources

Linguistic Context Recodes Visual Representations in Vision-Language Models

Goal-directed visual processing is a hallmark of human visual intelligence, resulting in representations that support downstream tasks such as categorization or search. Though vision-language models (VLMs) are often faced with these same tasks, their ability to recode visual representations when presented with goal-directed language remains poorly characterized. Indeed, prior work largely treats visual representations in VLMs as static repositories of visual information that are manipulated by language representations. In the present work, we provide evidence for two concrete instances of language-induced recoding of visual representations. First, we identify an abstract reference representation that denotes which objects are goal-relevant under a natural language prompt. We extract contrastive steering vectors corresponding to this reference representation and demonstrate that they are causally implicated in model predictions. These reference representations are abstract in that they generalize to different objects, different task contexts, and even from synthetic to naturalistic images. Second, we demonstrate language-induced attribute modulation: later layers selectively amplify goal-relevant attributes in visual representations of objects. We demonstrate this phenomenon across a range of different prompts. Finally, we provide a causal intervention that demonstrates that attribute modulation mediates a VLM's response distribution. Together, our results support a more dynamic account of cross-modality processing in VLMs -- rather than vision tokens serving as static repositories of information, they are modulated to support queries articulated in language.

cs.AI

Super-Resolution Posterior Ocular Microvascular Imaging Using 3-D Ultrasound Localization Microscopy With a 32X32 Matrix Array

The purpose of this study is to enable in-vivo three-dimensional (3-D) ultrasound localization microscopy (ULM) of posterior ocular microvasculature using a 256-channel system and a 1024-element matrix array, and to overcome limitations of restricted transmit angles, sound speed mismatch caused by the crystalline lens and surrounding tissues, and the low signal-to-noise ratio (SNR) of microbubble signals. To address phase distortions from the crystalline lens, which has a higher speed of sound (SOS) than surrounding tissues, a region-dependent SOS beamforming approach was implemented to improve microbubble resolution. A 4-D non-local means filter was subsequently applied to suppress background noise and enhance microbubble contrast. The proposed method improved localization accuracy and image quality, achieving a spatial resolution of 63 um, while Fourier shell correlation (1/2-bit threshold) confirmed a global resolution of approximately 59 um. Higher mean normalized cross-correlation coefficients between the microbubbles and the system point-spread function, obtained with the proposed method (approximately 0.67), compared with those without the proposed method (approximately 0.60), indicate enhanced microbubble signal quality. Furthermore, the 3-D bi-directional vessel density and flow-velocity maps were reconstructed, capturing detailed choroidal vascular and hemodynamic patterns. These results demonstrate that region-dependent SOS beamforming combined with spatiotemporal denoising enables high-resolution posterior ocular ULM and provides a practical pathway toward quantitative 3-D assessment of retinal and choroidal microvasculature for potential clinical use.

physics.med-ph