arXiv Science⌕ Search

arXiv · 2609.30647

Conditional Predictive Sufficient Statistics for Visual Representation Learning

Abstract

A useful visual representation is a statistic of the observed past that retains the latent factors shared with the future and discards patch-private noise. We formalize this requirement as a conditional predictive sufficient statistic (CPSS). Under a shared-factor model of image patches, the mutual information between the past and the next patch equals the information the past carries about the shared factor, up to a remainder that the next patch itself fails to reveal. Predicting the next patch embedding with a cosine loss is maximum likelihood for a von Mises-Fisher model of that embedding's direction, and is therefore a tractable surrogate for the predictive information. The same population loss is also minimized by a constant embedding, so stop-gradient does not by itself select the sufficient statistic; it only blocks the symmetric gradient that implements the constant solution in one step. The regression target is a shallow embedding, which forces the network output back into that shallow range and leaves the sufficient statistic in intermediate blocks. Small causal Transformers on MNIST and CIFAR-10 are used as diagnostics, not as a leaderboard. On MNIST the future shift and the stop-gradient move probe accuracy by tens of points, and the CPSS readout peaks before the output. On CIFAR-10, with the same short budget and no augmentation, every objective lands near a linear classifier on pixels. What still matches the derivation is the geometry: the CPSS output is a worse readout than its best intermediate block, next-pixel regression does not pay that penalty, and removing the stop-gradient collapses the effective rank of the embedding even when the pretext loss looks perfect.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yuzhou Hong. 2026-09-28. Conditional Predictive Sufficient Statistics for Visual Representation Learning. https://arxiv.org/abs/2609.30647

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Event-based Scene Synthesis via Inter-Frame Residual Alignment

Event-based scene synthesis reconstructs target RGB frames from sparse image observations and asynchronous event streams, encompassing both video frame prediction and interpolation. Existing event-based synthesis methods commonly estimate optical flow to warp the observed frames toward the target time, but are vulnerable to inaccurate flow under large motion and occlusion and often rely on flow supervision or pretrained estimators. In this work, we propose EvFRA, an Event-based scene synthesis framework based on inter-Frame Residual Alignment. We identify a structural correspondence between event measurements and frame-to-frame scene changes, and exploit this correspondence for target frame synthesis. Our training pipeline consists of two stages: 1) an Event-to-Residual Alignment Variational Autoencoder (ER-VAE) aligns the event frame captured between the anchor and target frames with the corresponding inter-frame residual, and 2) a ControlNet-conditioned diffusion model is fine-tuned to denoise the residual latent using event data. Our method outperforms state-of-the-art methods by up to 2.61 dB and 1.85 dB in PSNR for frame prediction and interpolation, respectively, with consistent SSIM improvements. Code is available at https://github.com/jiyun-kong/EvFRA.

cs.CV↗

From Concept Erasure to Style Purification: Contrastive Eigenbases for Artist Style Protection

Text-to-image diffusion models can reproduce specific artists visual styles at extremely low cost, raising copyright and deployment safety concerns about unauthorized style mimicry. Existing model-side protection methods generally follow ordinary concept erasure, emphasizing aggressive deletion or redirection of target styles. However, our causal intervention analysis shows that the central issue is not insufficient erasure strength, but a mismatch between artist styles and this paradigm: unlike ordinary object concepts, artist styles do not form compact, localized editable semantic units. Consequently, sparse editing and fixed retain lists struggle to suppress target styles while preserving generation utility. We therefore reformulate artist style protection as style purification, suppressing target style expression during inference while preserving the requested content and visual structure. We propose CAPE (Contrastive Artist Style Purification with Eigenbases), a training-free framework against artist style mimicry. CAPE constructs contrastive triplets around the target request and formulates style direction estimation as a generalized eigenvalue problem, capturing style-related directions that remain stable across content variations and are less affected by shared content. During inference, CAPE employs the Adaptive Suppression Controller to assign suppression strengths to different tokens based on Q, K, and V responses, and performs target style suppression on the K and V paths of self-attention. Experimental results show that CAPE effectively weakens target artist characteristics, including brushstrokes, textures, and local color processing, while better preserving major semantic entities, scene composition, and visual structures.

cs.CV↗

Evaluating Generative Models via One-Dimensional Code Distributions

Most evaluations of generative models rely on feature-distribution metrics such as FID, which operate on continuous recognition features that are explicitly trained to be invariant to appearance variations, and thus discard cues critical for perceptual quality. We instead evaluate models in the space of discrete visual tokens, where modern 1D image tokenizers compactly encode both semantic and perceptual information and quality manifests as predictable token statistics. We introduce Codebook Histogram Distance (CHD), a training-free distribution metric in token space, and Code Mixture Model Score (CMMS), a no-reference quality metric learned from synthetic degradations of token sequences. To stress-test metrics under broad distribution shifts, we further propose VisForm, a benchmark of 210K images spanning 62 visual forms and 12 generative models with expert annotations. Across AGIQA, HPDv2/3, and VisForm, our token-based metrics achieve state-of-the-art correlation with human judgments. We will release all code and datasets to facilitate future research, with the code publicly available at https://github.com/zexiJia/1d-Distance.

cs.CV↗