arXiv Science⌕ Search

arXiv subjects

Chuanzhi Zhou

Publications and source records attributed to Chuanzhi Zhou.

2 recordsLinked to original sources

Octree-based Video Representation

Video models commonly use uniform grids even though visual complexity varies substantially across space and time. We introduce OctVideo, which approximates a video clip with an octree. This hierarchy recursively partitions a spatio-temporal volume into eight subvolumes, so that smooth regions remain coarse while detailed regions receive finer cells. Each leaf stores local RGB values and spatio-temporal gradients, supplemented by a lightweight learned residual. For reconstruction, a Conv1D VAE maps the serialized cells to a regular latent grid and selectively refines details during decoding. Our VAE achieves 36.12 dB PSNR with 38.2M parameters and 189.4 GFLOPs per clip on Kinetics-400 (K400). It also generalizes zero-shot to the high-resolution Densely Annotated VIdeo Segmentation (DAVIS) 2016 dataset with reconstruction quality comparable to the best evaluated models. On both datasets, it requires the fewest model FLOPs and achieves the fastest encoding and decoding among the evaluated models. OctVideo also supports video understanding, achieving competitive recognition performance with few input tokens when trained from scratch. By exploiting the redundancy already present in video signals and efficiently processing sparse structures, OctVideo provides an efficient representation for video.

cs.CV↗

Geometry-Preserving Blind Watermarking for Raw 3D Point Clouds

Raw 3D point clouds are a core geometric representation. Establishing their ownership is challenging because point sets are irregular, unstructured, and frequently altered by resampling and geometric preprocessing. We present a blind watermarking framework that operates directly on xyz coordinates and supports both object-level shapes and scene-scale scans. At verification time, the embedded message is recovered from the observed point cloud alone, without access to the original point cloud, color, normals, or mesh connectivity. The method jointly learns watermark embedding and extraction through a feed-forward octree-based architecture, enabling efficient multi-scale geometric reasoning on large point sets. During training, a stochastic transformation layer exposes the decoder to common geometric perturbations, while progressive pose alignment improves robustness to pose changes. Experiments on object-level and scene-level benchmarks demonstrate reliable message recovery under common geometric processing while maintaining low geometric distortion. Qualitative comparisons further show that the learned perturbations are less visually conspicuous and less spatially structured than those of handcrafted alternatives.

cs.CV↗