arXiv ScienceSearch

arXiv subjects

Zhihao Xie

Publications and source records attributed to Zhihao Xie.

5 recordsLinked to original sources

SetEasy: A Multi-Modal Classroom Engagement Assessment and Seating Optimization Framework

SetEasy optimizes classroom engagement in fixed seating grids. It fuses multimodal sensing (wristband physiology, 4K video, environmental data) and trains a v-Gage model grounded in a revised ISEQ. Each week, two-week engagement forecasts are mapped to a student-seat utility matrix, and CP-SAT generates seating plans under visual-access and social-dynamics constraints. In a four-week deployment (23 students, 331 classes), v-Gage converged across affective, behavioral, cognitive, and overall dimensions, cutting RMSE from 0.75 to 0.53. Optimization raised mean engagement from 0.30 to 0.70, with over two-thirds of seats reaching high engagement and back-row low-activity patterns markedly reduced. These results show that, without hardware changes, interpretable, data-driven seating strategies can substantially enhance engagement. The multimodal "assessment + optimization" paradigm offers a transferable, sustainable path to culturally responsive, differentiated spatial design amid global homogenization.

cs.AI

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

Video generation models typically rely on 3D-VAEs trained for pixel-level reconstruction, whose latent spaces may underrepresent semantic structure. We introduce VideoRAE, a representation autoencoder that converts features from a frozen video foundation model into compact, reconstruction-capable latents for video generation. A lightweight 1D self-attention projector compresses multi-scale hierarchical features, producing continuous latents for Diffusion Transformers and discrete tokens for autoregressive models through multi-codebook high-dimensional quantization. During decoding, a local-global representation alignment objective transfers semantic structure from the frozen encoder and removes the need for KL regularization. Comprehensive experiments show that VideoRAE achieves strong reconstruction in both continuous and discrete regimes. On UCF-101, autoregressive and diffusion generators built on VideoRAE achieve class-conditional gFVD scores of 40 and 93, respectively, while converging approximately five times faster than autoencoder baselines. In controlled 2B-parameter text-to-video experiments, replacing LTX-VAE with VideoRAE accelerates convergence and consistently improves VBench performance. These results establish frozen video foundation representations as compact, versatile, and generation-friendly video latents. Code and models are available at https://zhxie0117.github.io/VideoRAE/.

cs.CV

CommuniWave:A Machine Learning Model for Quantifying the Degree of Temporary Informal Behavior in Urban Communities

For urban managers and designers, improving the functional attributes of urban communities to enhance territorial resilience in the face of complexity and uncertainty is crucial. Currently, community planning often follows a top-down approach and lacks effective metrics to quantify informal behaviors of residents, leading to frequent conflicts with original plans. This study introduces CommuniWave, a machine learning model designed to efficiently detect and quantify the Degree of Informal Behavior (DIB) in urban communities. The model integrates a Behavior Capture Net (BCN) based on mmaction2, a self-developed YOLOv10 model (YLX), and a Behavior Eval Model (BEM) using random forest. Ultimately, by generating DIB fluctuation charts from street videos, the model facilitates dynamic monitoring, supporting urban managers in making refined decisions to enhance the overall resilience of communities.

cs.AI

Parallel distributed quantum gates for dual-species quantum emitters

We propose a parallel protocol for implementing distributed nonlocal quantum gates between spatially separated stationary qubits encoded in dual-species quantum emitters (i.e., color-center and superconducting qubits). By utilizing entangled photon pairs with distinct frequencies as a quantum data bus, our approach connects spatially separated devices without requiring quantum frequency conversion or preshared entanglement, while maintaining an always-ready and resource-efficient property for distributed quantum computing and networks. Furthermore, we demonstrate the feasibility of implementing parallel distributed nonlocal quantum gates on multiple pairs of spatially separated qubits using a single high-dimensional entangled photon pair, which directly benefits from the enhanced quantum capacity provided by optical qudit encoding. Our protocol establishes a scalable and practically implementable framework for distributed quantum networks, potentially enabling the development of future large-scale quantum computing architectures.

quant-ph

TokBench: Evaluating Your Visual Tokenizer before Visual Generation

In this work, we reveal the limitations of visual tokenizers and VAEs in preserving fine-grained features, and propose a benchmark to evaluate reconstruction performance for two challenging visual contents: text and face. Visual tokenizers and VAEs have significantly advanced visual generation and multimodal modeling by providing more efficient compressed or quantized image representations. However, while helping production models reduce computational burdens, the information loss from image compression fundamentally limits the upper bound of visual generation quality. To evaluate this upper bound, we focus on assessing reconstructed text and facial features since they typically: 1) exist at smaller scales, 2) contain dense and rich textures, 3) are prone to collapse, and 4) are highly sensitive to human vision. We first collect and curate a diverse set of clear text and face images from existing datasets. Unlike approaches using VLM models, we employ established OCR and face recognition models for evaluation, ensuring accuracy while maintaining an exceptionally lightweight assessment process requiring just 2GB memory and 4 minutes to complete. Using our benchmark, we analyze text and face reconstruction quality across various scales for different image tokenizers and VAEs. Our results show modern visual tokenizers still struggle to preserve fine-grained features, especially at smaller scales. We further extend this evaluation framework to video, conducting comprehensive analysis of video tokenizers. Additionally, we demonstrate that traditional metrics fail to accurately reflect reconstruction performance for faces and text, while our proposed metrics serve as an effective complement.

cs.CV